Tutorials › AP Statistics › Mixed Practice: Describing Distributions in Context

Describing quantitative distributions · Tutorial 100 of 1000

Mixed Practice: Describing Distributions in Context

Learn to combine a histogram with summary statistics to write a complete, evidence-based description of a quantitative distribution.

Beginner 9 min read

What You'll Learn

  • Describe a distribution using evidence from both its histogram and summary statistics.
  • Choose measures of center and spread that suit the distribution’s shape and unusual values.
  • Check whether a reported minimum or maximum is flagged by the 1.5 IQR rule.
  • Compare two groups using histograms with common bins and matching summary statistics.
  • Write concise, contextual responses that respect what the display can and cannot show.

Combine the Graph and the Numbers

In Interpreting a Distribution to Answer a Question, you practiced matching a question to a feature that a graph can show. This tutorial brings a histogram and summary statistics together, as an exam-style free-response question might. The histogram helps you describe the overall pattern; the numerical summaries give precise information about center and spread.

The two sources of evidence should support one connected description. For example, a histogram with one central peak and similar-looking tails may support using the mean and standard deviation. A long tail or an unusually distant observation may make the median and IQR more useful. As in the earlier tutorials on the SOCS framework and choosing measures, describe shape and check for unusual values before deciding which summaries to emphasize.

Key idea: Treat the histogram and summary statistics as complementary evidence. Use the histogram for overall shape and concentrations, and use the numerical summaries for precise center and spread. Check that your claims fit both sources.

A histogram groups observations into intervals, so it does not show the exact values within each bin. Summary statistics can report exact quantities such as a mean or median, but a short list of statistics does not reveal every peak, gap, or concentration. A strong response uses each source for what it preserves, names the group and variable, includes units, and answers the question directly.

A Routine for Exam-Style Responses

A useful new technique is to make an evidence check as you draft: for each claim, identify whether the histogram, a summary statistic, or both support it. This keeps a numerical summary from replacing the description and prevents a visual impression from becoming an unsupported claim.

1
Identify the group and variable.
Read the prompt and axes carefully. Record the units, sample size, and what the observations represent.
2
Describe the histogram.
Use its bars to report the overall shape and the main concentration or peak. Mention a gap or unusual feature only if the graph supports it.
3
Check the summaries against the graph.
Choose a suitable center-and-spread pair. If quartiles and extremes are provided, use the 1.5 IQR rule when relevant. Do not assume that a histogram bin identifies an exact value.
4
Write a connected answer.
State the shape and any supported unusual features, then interpret suitable measures of center and spread in context. When comparing groups, use the same features and units for both.

The numbers below are invented for practice. In each example, the stated summaries are treated as calculator output from the same observations shown in the histogram. The task is not to calculate every summary from the raw data; it is to interpret the evidence and explain what it shows.

Worked Examples: Describe and Compare

Worked Example: Study Time for One Group

A fictional survey records how many hours 20 students spent studying on a particular evening. The frequency histogram has equal-width bins: 0 to less than 1 hour, 2 students; 1 to less than 2, 5; 2 to less than 3, 8; 3 to less than 4, 4; and 4 to less than 5, 1. The summaries are \(\bar{x}=2.55\) hours, \(s=1.00\) hour, median \(=2.60\) hours, \(Q_1=1.70\) hours, \(Q_3=3.20\) hours, minimum \(=0.40\) hour, and maximum \(=4.40\) hours. Describe the distribution, including a suitable center and spread.

Solution—State. We need to describe the evening study times of these 20 fictional students, using the histogram and the supplied summaries.

Plan. The histogram has one main peak, so first describe its overall shape. Then check the extremes using the 1.5 IQR rule. If there are no flagged extremes and the distribution is reasonably balanced, the mean and standard deviation are a suitable pair to report.

Do. The frequencies add to the stated sample size: \(2+5+8+4+1=20\). The histogram is unimodal, with its highest bar from 2 to less than 3 hours. Counts rise toward this bin and then fall; the pattern is roughly balanced, with no strong skew apparent from the bins.

The IQR is:

$$ \text{IQR}=Q_3-Q_1=3.20-1.70=1.50\text{ hours} $$

The lower and upper fences are:

$$ 1.70-1.5(1.50)=-0.55\text{ hours} \qquad 3.20+1.5(1.50)=5.45\text{ hours} $$

The minimum, 0.40 hour, and maximum, 4.40 hours, both fall within these fences. Thus, neither extreme is flagged by the 1.5 IQR rule. The mean and standard deviation are reasonable summaries for this roughly balanced distribution: the mean is 2.55 hours, and the standard deviation is 1.00 hour.

Conclude in context. “For these 20 fictional students, evening study time has one main concentration, in the 2-to-less-than-3-hour interval, and is roughly balanced with no extreme value flagged by the 1.5 IQR rule. The mean study time is 2.55 hours, and the standard deviation is 1.00 hour, indicating a typical distance of about 1 hour from the mean.”

The histogram supports describing a peak interval, not claiming that exactly 2.55 hours was the most common amount. The mean and standard deviation summarize center and spread; they do not replace the description of the histogram’s pattern.

Worked Example: Right-Skewed Delivery Delays

A fictional sample of 28 deliveries records delay in minutes. A histogram has these counts: 0 to less than 5 minutes, 12; 5 to less than 10, 8; 10 to less than 15, 4; 15 to less than 20, 2; 20 to less than 25, 1; and 25 to less than 30, 1. The summaries are mean \(=7.1\) minutes, standard deviation \(=7.0\) minutes, median \(=5.0\) minutes, \(Q_1=2.5\) minutes, \(Q_3=10.0\) minutes, minimum \(=0\) minutes, and maximum \(=29\) minutes. Describe the distribution and select appropriate summaries.

Solution. The counts sum to \(12+8+4+2+1+1=28\), matching the sample size. The histogram has its greatest concentration in the first two bins, and the bars extend across successively larger delay intervals, with a sparse tail to the right. This is a right-skewed distribution.

Check the reported maximum with the IQR rule:

$$ \text{IQR}=10.0-2.5=7.5\text{ minutes} $$
$$ \text{Upper fence}=10.0+1.5(7.5)=21.25\text{ minutes} $$

The maximum delay of 29 minutes is above the upper fence of 21.25 minutes, so it is flagged as an outlier by the 1.5 IQR rule. The minimum of 0 minutes is above the lower fence, \(2.5-1.5(7.5)=-8.75\) minutes, and is not flagged. A flag identifies a value for further attention; it does not explain why the delay occurred.

Because the distribution is right-skewed and contains a flagged high value, the median and IQR are more appropriate than the mean and standard deviation for describing a typical delay and the spread of the middle half. The median delay is 5.0 minutes. The IQR is 7.5 minutes, meaning the middle 50% of delays span 7.5 minutes, from \(Q_1=2.5\) to \(Q_3=10.0\) minutes.

Answer in context. “For these 28 fictional deliveries, delays are concentrated below 10 minutes, with a long tail toward larger delays. The median delay is 5.0 minutes, and the middle half spans 2.5 to 10.0 minutes (an IQR of 7.5 minutes). The maximum delay of 29 minutes is flagged by the 1.5 IQR rule.”

The mean of 7.1 minutes is greater than the median of 5.0 minutes, which is consistent with the right tail pulling the mean upward. Reporting the mean is not mathematically wrong, but the median and IQR better represent the center and spread for this skewed distribution.

Worked Example: Comparing Two Check-In Groups

A fictional clinic compares check-in wait times for 24 morning appointments and 24 afternoon appointments. Both histograms use the same bins, in minutes. Morning counts are 1, 4, 11, 6, and 2; afternoon counts are 7, 8, 5, 3, and 1. The bins are 0 to less than 5, 5 to less than 10, 10 to less than 15, 15 to less than 20, and 20 to less than 25 minutes. For morning appointments, the mean is 11.2 minutes, median 11.4 minutes, \(Q_1=8.8\) minutes, \(Q_3=14.6\) minutes, and standard deviation 4.5 minutes. For afternoon appointments, the mean is 7.8 minutes, median 6.1 minutes, \(Q_1=3.0\) minutes, \(Q_3=10.2\) minutes, and standard deviation 5.4 minutes. Compare the distributions.

Solution. First check that each histogram represents 24 appointments. Morning counts total \(1+4+11+6+2=24\), and afternoon counts total \(7+8+5+3+1=24\). Since both histograms use the same bins, their bar patterns can be compared directly.

The morning distribution has one main concentration in the 10-to-less-than-15-minute interval. The afternoon distribution has its largest counts in the first two intervals, below 10 minutes, and extends toward longer waits. The afternoon mean, 7.8 minutes, is above its median, 6.1 minutes, consistent with a longer right tail. The histograms suggest that morning waits are centered at higher values overall.

Compare spread using the IQRs:

$$ \begin{aligned} \text{Morning IQR}&=14.6-8.8=5.8\text{ minutes}\\ \text{Afternoon IQR}&=10.2-3.0=7.2\text{ minutes} \end{aligned} $$

The afternoon IQR is larger, so the middle half of afternoon waits covers a wider interval. The standard deviations point in the same direction: 5.4 minutes for afternoon appointments compared with 4.5 minutes for morning appointments. The afternoon distribution therefore shows more spread by both measures provided. The afternoon median is lower, 6.1 minutes compared with 11.4 minutes in the morning, but its longer right tail means not every afternoon wait is short.

Answer in context. “In these fictional appointments, morning waits are concentrated most heavily from 10 to less than 15 minutes, while afternoon waits are concentrated below 10 minutes and extend toward larger values. The morning median is 11.4 minutes, compared with 6.1 minutes in the afternoon. Afternoon waits have greater spread in the middle half (IQR 7.2 versus 5.8 minutes) and a larger standard deviation (5.4 versus 4.5 minutes).”

This comparison describes the observed groups. It does not establish why their wait times differ or prove that the same pattern holds for all clinic appointments.

Common Mistakes and AP Exam Tips

  • Listing statistics without describing the histogram. A response that reports a median and IQR but says nothing about shape leaves out evidence the histogram provides. Name the main pattern or concentration.
  • Using a bin as if it were an exact value. A peak in the 2-to-less-than-3-hour bin supports calling that the peak interval. It does not show that 2.5 hours is the most frequent exact value.
  • Choosing a center-and-spread pair that does not fit the distribution. For skew or a flagged outlier, the median and IQR are generally more informative. For an approximately symmetric distribution without strong outliers, the mean and standard deviation are often suitable.
  • Calling an observation an outlier without a basis. If using the 1.5 IQR rule, show the IQR and relevant fence, then state whether the observation lies beyond it. A flagged value is not proof of a mistake or an explanation for its cause.
  • Comparing groups with vague language. Replace “Group A is higher” with a contextual statement specifying what is higher, such as “the median morning wait is 11.4 minutes, compared with 6.1 minutes in the afternoon.” Use matching measures and units.
  • Making a population claim from the displayed sample. Describe the observations shown unless the question provides a basis for a broader claim. A descriptive comparison alone does not explain a difference or establish that it applies to everyone.

For full credit, connect claims to evidence. For instance, “The delays are right-skewed, with most below 10 minutes and a tail toward larger values; the median is 5.0 minutes and the IQR is 7.5 minutes” links visible shape to appropriate summaries and gives units. If the question asks for a comparison, describe each group on the same features before stating the difference.

Key takeaway: Read the histogram for shape and concentrations, check unusual values when the summaries allow it, and choose measures of center and spread that fit the distribution. Link each claim to evidence and describe the group and variable in context.

Check Your Understanding

Use the histogram information and summaries in each question to plan a careful response.

  1. A histogram of 18 fictional quiz-completion times has a single peak and a roughly balanced pattern. The mean is 14.2 minutes and the standard deviation is 2.1 minutes. Which center-and-spread pair is reasonable to report, and what context should your interpretation include?
  2. A right-skewed histogram of household water use has median 310 liters, \(Q_1=250\) liters, and \(Q_3=390\) liters. Find the IQR and state why it may be more informative than the mean and standard deviation.
  3. A histogram’s tallest bar covers 8 to less than 12 kilograms. What is the most specific claim about the peak that this display supports?
  4. For a group, \(Q_1=6\) minutes and \(Q_3=14\) minutes. Find the upper fence using the 1.5 IQR rule and decide whether a maximum of 28 minutes is flagged.
  5. Two groups have histograms with different bin widths. What should you check before comparing their bar heights, and what other information could support the comparison?