Start With the Question, Then Choose the Evidence
In “Mapping Course Topics to Exam Questions,” you practiced identifying the statistical task in a prompt. This tutorial takes the next step: once you know what the question asks, how do you choose a graph and a summary that answer it? The key is to match the evidence to both the type of data and the claim or feature of interest.
A graph helps reveal a distribution or relationship. A summary statistic compresses part of that information into a number. Neither is automatically useful just because it is available. For example, a mean does not show whether a distribution is skewed, and a correlation does not describe a curved relationship well. Begin by identifying the variables and asking what the question wants you to compare or describe.
A practical routine is to work through four decisions:
Decide whether the data are categorical or quantitative, and whether the question concerns one variable, a relationship, or a comparison between groups.
Is it asking for a typical value, variability, shape, unusual observations, association, or a difference between groups?
Select a graph that makes the relevant pattern visible and a statistic that measures the requested feature.
Say what the graph or statistic reveals, use the variable’s context and units, and note an important limitation if it affects the answer.
A Question-First Guide to Graphs and Summaries
The data structure narrows the graph choices. The wording of the question narrows them further. This guide gives common matches; it is not a requirement to report every statistic in a row.
| Data and question | Useful graph | Useful summary |
|---|---|---|
| One quantitative variable: shape, gaps, or individual values | Dotplot or stemplot for a modest number of observations; histogram for a larger set | Describe shape and unusual features; use context to explain what stands out |
| One quantitative variable: typical value and spread | Dotplot, histogram, or boxplot, as appropriate | Mean and standard deviation for a reasonably symmetric distribution without strong outliers; median and IQR when skew or outliers make resistant summaries more appropriate |
| One categorical variable: counts or proportions in categories | Bar chart | Category counts or proportions, depending on the question |
| Two quantitative variables: relationship | Scatterplot | Describe form, direction, and strength; use \(r\) to summarize direction and strength of a linear association |
| Quantitative response compared across groups | Side-by-side boxplots or comparative histograms | Compare centers and spreads, choosing summaries suited to the distributions |
| Categorical response compared across groups | Side-by-side bar charts or segmented bar charts | Compare within-group proportions for each response category |
The choice between mean and median, and between standard deviation and IQR, depends on what the distribution looks like and what you need to describe. As in earlier tutorials on exploring one-variable data, the mean and standard deviation are more affected by extreme values than the median and IQR. Graphs help you judge whether that difference matters here.
For comparisons, keep the groups visible. A single statistic for all observations can conceal differences between groups. Similarly, when group sizes differ, comparing raw category counts may be misleading; within-group proportions are often more informative. For two quantitative variables, a scatterplot keeps the paired values together, while separate lists of means would discard information about their relationship.
Worked Examples: Match the Summary to the Question
Worked Example: Typical Commute Time and a Long Trip
Situation. A fictional student group records nine one-way commute times, in minutes: 12, 14, 15, 16, 18, 19, 20, 22, and 41. The prompt asks for a typical commute time and whether the data contain an unusual feature.
Identify the task. There is one quantitative variable, commute time. The question asks about center and an unusual observation. Because a dotplot can show each of these nine values, it is a useful display for seeing both the cluster and the long commute. A boxplot can summarize the center and spread compactly.
Calculate summaries that fit the distribution. The ordered values have a median of 18 minutes. Excluding the median, the lower half is 12, 14, 15, 16, so \(Q_1=(14+15)/2=14.5\) minutes. The upper half is 19, 20, 22, 41, so \(Q_3=(20+22)/2=21\) minutes. Therefore:
The mean provides a useful contrast:
Interpret the match. The median commute was 18 minutes, and the middle half of the recorded times spans 6.5 minutes. The 41-minute trip lies well above the other values. The mean is pulled upward by that long trip, so the median and IQR better summarize a typical time and the spread of the bulk of these observations. A dotplot makes the isolated high value especially clear.
Why not give only the mean? The mean answers a different question: it is the arithmetic average, and here it is affected by the unusually long trip. A full-credit justification links the choice of median and IQR to the visible high value and the desire to describe typical commutes, rather than simply stating that these statistics are “better.”
Worked Example: Compare the Spread of Two Wait-Time Groups
Situation. In a fictional service-center exercise, students record these wait times, in minutes, for two separate groups of eight visitors:
| Group A | Group B |
|---|---|
| 5, 6, 6, 7, 7, 8, 8, 9 | 4, 5, 5, 6, 8, 9, 10, 11 |
The question asks whether the groups have similar typical wait times and which group has more variability. These are two distributions of a quantitative variable, so side-by-side boxplots can make their centers and spreads easy to compare. A comparison should use the same type of summary for each group.
Find the centers and IQRs. For Group A, the median is \((7+7)/2=7\) minutes. Its lower half is 5, 6, 6, 7, giving \(Q_1=(6+6)/2=6\), and its upper half is 7, 8, 8, 9, giving \(Q_3=(8+8)/2=8\). Thus its IQR is \(8-6=2\) minutes.
For Group B, the median is \((6+8)/2=7\) minutes. Its lower half is 4, 5, 5, 6, giving \(Q_1=(5+5)/2=5\), and its upper half is 8, 9, 10, 11, giving \(Q_3=(9+10)/2=9.5\). Thus its IQR is \(9.5-5=4.5\) minutes.
Compare in context. Both groups have a median wait of 7 minutes, so their centers are the same by this measure. Group B has the larger IQR, 4.5 minutes compared with 2 minutes, indicating more spread in the middle half of its wait times. Side-by-side boxplots would show this difference in spread while making the matching medians visible.
State the limitation. An IQR comparison describes the middle half of each group; it does not show every feature of the distributions. If the question also asks about shape or individual unusual times, inspect a graph that displays more detail, such as comparative dotplots or histograms.
Worked Example: Detect an Association That Correlation Misses
Situation. A fictional energy model records a temperature setting’s deviation from a target, \(x\), in degrees, and an energy-use index, \(y\), for five settings:
| Deviation \(x\) (degrees) | Energy-use index \(y\) |
|---|---|
| -2 | 4 |
| -1 | 1 |
| 0 | 0 |
| 1 | 1 |
| 2 | 4 |
The prompt asks whether energy use is associated with the setting. Both variables are quantitative, so the first choice should be a scatterplot. The points form a U-shaped pattern: the energy-use index is lower near the target and higher at deviations on either side. This is an association, even though it is not linear.
Check what \(r\) says. The means are \(\bar{x}=0\) and \(\bar{y}=(4+1+0+1+4)/5=2\). The sum of the products of deviations is:
The sums of squared deviations are \(\sum (x-\bar{x})^2=10\) and \(\sum (y-\bar{y})^2=14\). Therefore:
Interpret the evidence. The correlation is zero, so it indicates no linear association in these data. It does not mean the variables are unrelated: the scatterplot reveals a clear curved association. The graph answers the broader question about association; \(r\) alone would miss its form. As emphasized in “Choosing Numbers to Support a Linear Model,” use \(r\) to summarize direction and strength when a linear relationship is relevant, and inspect the scatterplot rather than treating \(r\) as a complete description.
Worked Example: Compare Categorical Responses Fairly
Situation. A fictional survey asks students in two grades whether they usually bring lunch from home. Among 40 ninth-graders, 24 answer yes. Among 50 twelfth-graders, 20 answer yes. The question asks whether the response patterns differ by grade.
Match the display to the variables. Grade is categorical, and lunch response is categorical. A side-by-side bar chart or segmented bar chart can compare the response categories across grades. Since the group sizes differ, compare the proportions within each grade rather than just comparing the counts of “yes” responses.
Calculate the within-grade proportions. For ninth-graders, \(24/40=0.60\), or 60%, usually bring lunch from home, leaving \(16/40=0.40\), or 40%, who do not. For twelfth-graders, \(20/50=0.40\), or 40%, usually bring lunch from home, leaving \(30/50=0.60\), or 60%, who do not.
Interpret the comparison. In this survey, the proportion usually bringing lunch from home is 20 percentage points higher among ninth-graders than among twelfth-graders: \(60\%-40\%=20\) percentage points. A segmented bar chart using within-grade percentages would display the differing response patterns clearly. The data describe these surveyed students; the graph and proportions alone do not establish why the patterns differ or justify a broader population claim.
Common Mistakes and AP Exam Tips
- Choosing a statistic before looking at the distribution. A mean and standard deviation can be misleading summaries of a strongly skewed distribution or one with an influential extreme value. Examine a suitable graph, then justify summaries that fit the pattern.
- Treating “typical” as a command to use the mean. A typical value might be summarized by the mean or median. Explain which is more informative for the distribution and question at hand.
- Using a graph that hides the structure of the data. A histogram can show shape but not preserve the identity of each observation. A scatterplot is needed to display paired quantitative values; a bar chart is appropriate for categories.
- Using \(r\) as a complete measure of association. Correlation summarizes linear direction and strength. In the energy-use example, \(r=0\) despite a clear U-shaped pattern. Describe the scatterplot’s form as well.
- Comparing counts when group sizes differ. A larger group can have a larger count simply because it contains more people. When the question is about how response patterns differ, calculate and compare within-group proportions.
- Listing a graph or statistic without explaining why it fits. “Use a boxplot” is not a full justification. Say that side-by-side boxplots display the distributions’ centers and spreads for a quantitative response across groups.
- Confusing a descriptive difference with an explanation. A graph can show that groups differ in the observed data. It does not, by itself, establish a cause or explain the difference.
A strong AP response names the feature the prompt asks about, identifies evidence that directly displays or measures it, and interprets that evidence with the variable and units. If you choose a resistant summary, point to the skewness or unusual values that make it appropriate. If you compare groups, state which group has the larger or smaller value and specify whether you are comparing centers, spreads, counts, or proportions.
Check Your Understanding
For each prompt, identify a suitable graph and summary or description, and explain why that choice addresses the question.
- A set of 18 home-sale prices has a long right tail. The question asks for a typical price and the spread of the middle half. What summaries and graph would you choose?
- A school compares test scores, a quantitative variable, across three course sections. Which display could make the centers and spreads easy to compare?
- A researcher asks whether two quantitative measurements are associated. What graph should come first, and what feature does \(r\) summarize?
- Two groups with different sample sizes answer yes or no to a question. If the goal is to compare response patterns, should you compare counts or within-group proportions? Explain.
- A scatterplot shows a strong curved pattern, but \(r\) is close to zero. What conclusion about the variables’ association is justified?