Look at the Distributions Before Testing
A two-sample t test compares population means, but the test should not be the first thing you look at. Before calculating a test statistic, inspect the two groups’ distributions. Side-by-side boxplots make it easier to compare their centers and spreads and to notice features such as skewness or possible outliers.
As covered in “Conditions for a Two-Sample t Test,” the t procedure requires appropriate study design and data that support its use. A boxplot can help assess the shape of each group’s data, especially when a sample is small. It cannot show whether the groups were randomly sampled, whether observations are independent, or whether the samples were randomly assigned. Those questions come from the study design.
What a Boxplot Shows
A boxplot displays a compact summary of a quantitative distribution. The box runs from the first quartile \(Q_1\) to the third quartile \(Q_3\), with a line marking the median. The box’s length is the interquartile range, or IQR, which measures the spread of the middle half of the data. Whiskers extend beyond the box. In a modified boxplot, separate marks may identify possible outliers beyond the whiskers.
When comparing two groups, describe each distribution on its own before saying how the groups differ. A useful order is shape, center, spread. Shape includes whether the distribution appears roughly symmetric or skewed and whether possible outliers are marked. The median gives a measure of center, and the IQR or range gives a measure of spread. Use the same units for both groups.
Boxplots provide clues about shape, not a complete picture. A longer whisker on one side or a median off-center in the box may suggest skewness. A point marked beyond a whisker is a possible outlier under the graph’s convention. But a boxplot does not display every observation, so it may hide clusters, gaps, or other details visible in a dotplot or histogram. Also, a possible outlier is a reason to investigate the data and consider its effect—not an automatic reason to delete it.
Make the Comparison Fair
Side-by-side boxplots are most useful when both groups use the same scale and axis. If one graph uses a different scale, differences in the lengths of boxes or whiskers can be misleading. Label the groups, identify the response variable and units, and make clear which way larger values lie.
Compare the medians to describe the groups’ typical observed values, and compare the IQRs to describe the middle-half spreads. The range can help show the overall extent of the data, but it is sensitive to extreme values. If the boxplots show possible outliers, mention them and consider whether they could affect a mean-based analysis. A t procedure uses sample means and standard deviations; the median and IQR are visual summaries, not substitutes for those quantities.
Worked Examples
Worked Example: Similar Shapes, Different Centers
A school’s environmental club compares the number of minutes students spend watering a garden during a scheduled volunteer shift. The two independent groups are students using a drip system and students using a hose. Side-by-side boxplots, made on the same scale, have these five-number summaries:
| Group | Minimum | \(Q_1\) | Median | \(Q_3\) | Maximum |
|---|---|---|---|---|---|
| Drip system | 12 | 18 | 24 | 31 | 40 |
| Hose | 10 | 16 | 22 | 28 | 36 |
Describe shape, center, and spread before considering a test comparing the population mean watering times.
Describe shape: In the drip-system group, the median is 6 minutes above \(Q_1\) and 7 minutes below \(Q_3\); the whiskers extend 6 minutes below the box and 9 minutes above it. These distances are fairly balanced, so the plot does not show strong skewness. The hose group also looks roughly balanced: the median is 6 minutes above \(Q_1\) and 6 minutes below \(Q_3\), and its whiskers extend 6 minutes below and 8 minutes above the box. Neither summary shows a marked possible outlier.
Compare center: The drip-system sample has a median of 24 minutes, compared with 22 minutes for the hose sample. In these samples, the typical watering time is 2 minutes higher with the drip system, based on the medians.
Compare spread: The IQRs are
The ranges are \(40-12=28\) minutes for the drip group and \(36-10=26\) minutes for the hose group. The middle half and the overall range are slightly more spread out in the drip group. These are modest visual differences, not evidence by themselves that the population means differ.
Before inference: The plots show no obvious severe skewness or marked outliers, which is reassuring for a t procedure’s shape check. The graph does not tell us whether students were randomly sampled or whether their observations are independent; those details must be checked from how the data were collected.
Worked Example: A Possible Outlier and Right Skew
A product-design class compares the operating times, in hours, of two types of portable reading lamp. Each group has 12 lamps. The modified boxplots use the same axis. For Type A, the box extends from 5 to 8 hours, the median is 6 hours, and the whiskers run from 4 to 12 hours; a separate point appears at 20 hours. For Type B, the minimum is 7 hours, \(Q_1=8\), the median is 9, \(Q_3=10\), and the maximum is 11 hours, with no separate outlier mark.
Describe shape: Type A’s median is closer to the bottom of its box: it is 1 hour above \(Q_1\) and 2 hours below \(Q_3\). Its upper whisker is much longer than its lower whisker, and it has a high marked point. Together, these features suggest right skewness and identify 20 hours as a possible outlier. Type B looks comparatively balanced: its median is 1 hour from each quartile, and its whiskers are each 1 hour long.
Compare center: The sample median is 6 hours for Type A and 9 hours for Type B. Thus, the typical observed operating time, as measured by the median, is 3 hours higher for Type B.
Compare spread: Type A’s IQR is \(8-5=3\) hours. Type B’s IQR is \(10-8=2\) hours. For Type A, the high marked value also makes the total range \(20-4=16\) hours; Type B’s range is \(11-7=4\) hours. The IQR and especially the range show more spread in Type A. The point at 20 hours is far above the rest of that group’s plot, so it could have a substantial effect on Type A’s sample mean and standard deviation.
Before inference: With only 12 lamps per group, the pronounced right skew and possible outlier in Type A are important concerns for a two-sample t procedure. The plots do not support simply treating the group as approximately Normal. The marked value should be checked for a recording or measurement issue, but it should not be removed merely because it is unusual. The study design also needs to establish random selection or assignment and independent groups; the boxplots cannot establish either condition.
Worked Example: Overlapping Boxes Do Not Decide the Test
A town transit team records the waiting time, in minutes, for randomly selected riders at two stops during comparable weekday periods. Stop 1 has 36 sampled riders and Stop 2 has 40. The samples are independent, and the rider population at each stop is more than ten times the sample size. Their boxplots show these summaries:
| Stop | Minimum | \(Q_1\) | Median | \(Q_3\) | Maximum |
|---|---|---|---|---|---|
| Stop 1 | 3 | 6 | 9 | 13 | 19 |
| Stop 2 | 2 | 7 | 11 | 15 | 22 |
Describe shape: Both plots have somewhat longer upper than lower whiskers, so each may be mildly right-skewed. Neither has a marked outlier. A boxplot does not reveal every detail of the distributions, but these summaries do not suggest severe skewness or an extreme point.
Compare center and spread: Stop 2’s sample median is \(11-9=2\) minutes higher than Stop 1’s. The IQRs are \(13-6=7\) minutes at Stop 1 and \(15-7=8\) minutes at Stop 2. Their ranges are \(19-3=16\) minutes and \(22-2=20\) minutes, respectively. Stop 2’s observed waiting times are slightly more spread out by both measures.
Connect the plots to a possible test: Let \(\mu_1\) and \(\mu_2\) be the true mean waiting times for riders at Stops 1 and 2 during the specified weekday periods. The boxplots give a preliminary look at whether a two-sample t test might be reasonable; they do not supply the sample means or standard deviations needed to calculate the test statistic.
Check the conditions: The random samples support inference to riders in the stated populations, assuming the selection process was carried out as described. The samples are less than 10% of their respective rider populations, so the 10% condition is satisfied for sampling without replacement. Each rider is measured at one stop, and the problem states that the samples are independent, so there is no pairing between groups. Finally, both sample sizes are at least 30, and the boxplots show no severe skewness or marked outliers; this supports using a two-sample t procedure. The visual check does not replace the random-design and independence checks.
The higher sample median at Stop 2 is a descriptive observation, not a conclusion about the population means. A later test would use the sample means, standard deviations, sample sizes, and the conditions just reviewed to assess evidence about \(\mu_1-\mu_2\).
Common Mistakes and AP Exam Tips
- Calling the median the mean: The line inside the box marks the median. A boxplot does not display the sample mean, so describe the plotted center as a median unless the mean is separately provided.
- Reporting numbers without context: Instead of saying “the median is 11,” say “the sample median waiting time at Stop 2 is 11 minutes.” Include the group and units.
- Comparing plots on different scales: Use a common axis. Otherwise, apparent differences in box or whisker lengths may not reflect comparable amounts of spread.
- Claiming that a higher median proves a higher population mean: The plot describes the observed samples. A test is needed to assess evidence about population means, and its result depends on the sample means and standard errors—not just the medians.
- Treating a marked point as an error that must be deleted: A possible outlier should be investigated and its effect considered. Do not remove a valid observation solely because it is unusual.
- Using the graph to claim independence or random sampling: These are design conditions. Explain them from the way the study was carried out, not from the appearance of the boxplots.
- Assuming overlapping boxes mean no difference: Box overlap is not a decision rule for a two-sample t test. Describe the plots, then use the appropriate inference procedure to assess evidence about the means.
For a clear AP response, name the response variable and units, describe each group’s shape, compare medians and spreads, and mention possible outliers. Then connect the plot to the t procedure cautiously: say whether the observed shapes raise concerns, and check study design and independence separately. Avoid claiming that the plot proves a population difference or guarantees that the procedure is valid.
Check Your Understanding
Use the boxplot summaries and descriptions in this tutorial to answer each question.
- For the garden-watering data, calculate each group’s IQR and describe which sample has the larger one.
- In the lamp example, what plot features suggest that Type A may be right-skewed, and why does the small sample size make those features important before a t procedure?
- What does a median line inside a boxplot represent? Does it give the sample mean?
- In the transit example, name one condition supported by the study design and one feature assessed from the boxplots.
- Why is overlap between two boxes not, by itself, a conclusion about whether the population means differ?