Tutorials › AP Statistics › Mixed Practice: Interpreting Categorical Graphs

Graphs for categorical data · Tutorial 60 of 1000

Mixed Practice: Interpreting Categorical Graphs

Practice choosing a suitable categorical graph, interpreting what its values show, and identifying errors or unsupported claims.

Beginner 9 min read

What You'll Learn

  • Decide whether a question calls for a graph of one categorical variable or a comparison across groups.
  • Choose a display that makes the requested comparison clear.
  • Read graph values carefully and distinguish counts from percentages.
  • Identify misleading scales, missing labels, and claims that go beyond the display.
  • Write conclusions that use relevant evidence and stay within the scope of the data.

Mixed Practice: One Graph, Several Decisions

A categorical graph question may ask you to choose a display, read a value, compare groups, or explain why a graph or conclusion is misleading. Those tasks use many of the same ideas, but they do not always require the same graph or the same evidence. First work out what the question is asking; then inspect the display and support your answer with the relevant features or values.

In Choosing Between Bar Chart, Pie Chart, and Histogram, you learned to match a graph to the type of variable. For categorical data, bar charts and pie charts display one categorical distribution. When a question compares groups, side-by-side bars, segmented bars, and mosaic plots can show group differences in different ways. Earlier tutorials such as Bar Chart Versus Segmented Bar Chart: Picking the Right One explain when those displays are useful.

This tutorial brings those decisions together. The new habit to practice is to treat each graph question as a short sequence: identify the task, check what the graph encodes, and then state only what the display supports. A graph can be visually striking but still fail to answer the question if it shows counts where percentages are needed, or if its labels and scale are unclear.

Key idea: Before reading or judging a categorical graph, identify the variable, the groups (if any), and what the question asks you to compare. Then check whether the graph displays the right quantities for that task.

A Three-Part Routine for Exam Questions

Do not begin by describing every visible feature. Begin by deciding what kind of answer is needed. A question asking “Which response was most common?” calls for the category with the greatest count or proportion. A question asking “Which group had the greater share?” calls for within-group percentages. A question asking whether a display is misleading calls for a specific feature and an explanation of how it could affect a reader.

1
Identify the graph’s job.
Is the graph describing one categorical variable, comparing a response across groups, or showing both group size and response composition?
2
Check what is encoded.
Read the title, category labels, scale, units, and legend. Determine whether bar heights or segment sizes represent counts, proportions, or percentages, and identify the group represented by each bar or section.
3
Answer with matching evidence.
Use counts for a question about numbers of individuals, and within-group percentages for a question about how common a response is in each group. For a critique, name the problem and its effect.

This routine also helps you decide whether a graph is appropriate. A separated-bar display is appropriate for categories; a histogram is for quantitative data, as explained in Why Histograms and Bar Charts Are Different. If the question is about comparing two groups’ response distributions, a single bar chart of the combined responses may hide those group differences. A graph’s design should make the requested comparison visible rather than forcing the reader to reconstruct it.

When a graph has no printed data labels, read values from the scale and report estimates as approximate. When the display gives exact labels or a table of counts, use those values rather than guessing from bar heights. As in Explaining a Claim From a Categorical Graph, a strong answer names the category or groups and connects the evidence to the question.

Worked Example: Choose a Display for One Categorical Variable

Worked Example: Choose a Display for One Categorical Variable

A fictional community center asks 80 visitors which workshop they would most like to attend. The responses are composting, herb gardening, native plants, or another topic. The counts are 32, 24, 16, and 8, respectively. The center wants to show which topic was most popular and how the responses were divided. Which graph would you recommend, and what does it show?

Identify the task. There is one categorical variable: preferred workshop topic. The center wants to compare category sizes and summarize the whole distribution. A bar chart is a clear choice because the separated bars make it easy to compare the category counts. A pie chart could also show each topic as a part of the total, but it may be harder to compare similar-sized slices precisely.

Check the totals and convert if helpful. The counts add to the stated total:

$$ 32+24+16+8=80. $$

The relative frequencies and percentages are:

$$ \begin{aligned} \text{Composting: }&\frac{32}{80}=0.40=40\%,\\ \text{Herb gardening: }&\frac{24}{80}=0.30=30\%,\\ \text{Native plants: }&\frac{16}{80}=0.20=20\%,\\ \text{Another topic: }&\frac{8}{80}=0.10=10\%. \end{aligned} $$

The percentages add to \(100\%\), as expected for the full distribution. A relative-frequency bar chart would show the same pattern on a percentage scale instead of a count scale. The display choice depends on the intended message: counts show how many visitors gave each response, while percentages show each topic’s share of the 80 responses.

Answer in context. A bar chart of counts is a good choice for comparing how many visitors selected each topic. Composting is the most popular response in this sample, with 32 of the 80 visitors, or 40%. The display describes these surveyed visitors; it does not by itself establish which workshop would be most popular among all possible community-center visitors.

A complete exam response should give a reason for the choice, not just name a graph. For example, “Use a bar chart because the variable is categorical and the center wants to compare the response counts” ties the graph to both the variable and the task.

Worked Example: Read a Group Comparison Without Confusing Counts and Percentages

Worked Example: Read a Group Comparison Without Confusing Counts and Percentages

A fictional trail committee surveys people arriving for two types of cleanup events. Among 50 weekday-event participants, 30 say they would volunteer again. Among 100 weekend-event participants, 45 say they would volunteer again. A side-by-side bar chart displays the counts of “yes” and “no” responses for each event type. Which type of event has the greater proportion of participants who would volunteer again? Which has the greater count?

Read what the graph displays. The graph shows counts. That is enough to answer which event type has more surveyed “yes” responses, but not which event type has the larger proportion saying “yes,” because the group sizes differ.

Find each group’s percentage. For weekday events, 30 of 50 participants say yes:

$$ \frac{30}{50}=0.60=60\%. $$

For weekend events, the 45 yes responses are out of 100 participants:

$$ \frac{45}{100}=0.45=45\%. $$

The no counts are 20 for weekday events and 55 for weekend events. They agree with the totals: \(30+20=50\) and \(45+55=100\). The corresponding no percentages are \(20/50=40\%\) for weekday events and \(55/100=55\%\) for weekend events.

Answer both questions separately. Weekend events have the greater count of participants who would volunteer again: 45, compared with 30 for weekday events. Weekday events have the greater proportion saying yes: 60%, compared with 45% for weekend events. The difference in the yes percentages is:

$$ 60\%-45\%=15\text{ percentage points}. $$

A segmented bar chart would be a useful choice if the main goal were to compare the response distributions within the two event types, because each bar would represent \(100\%\) of its own group. A mosaic plot could also show the response distribution while representing that the weekend group is larger. The existing count chart is not automatically wrong; it answers a different question clearly. The key is to match the evidence to the question, as in Comparing Groups of Different Sizes With Percent Graphs.

Worked Example: Critique a Scale and a Claim

Worked Example: Critique a Scale and a Claim

A fictional neighborhood survey asks residents whether they support adding a protected bike lane. A bar chart reports 68% support among 50 Eastside respondents and 74% support among 50 Westside respondents. Its vertical axis begins at 60% rather than 0%, and the chart title says “Bike lane support across the city.” A student concludes, “Westside residents overwhelmingly support the plan, so the city should build it.” Identify two problems with the graph or claim and write a more careful conclusion.

Critique the scale. The bar chart’s vertical axis begins at 60%. In a bar chart, bar lengths encode values, so omitting the part of the scale below 60% makes the visible bars look much more different in height than their actual values. The support percentages differ by:

$$ 74\%-68\%=6\text{ percentage points}. $$

With a zero baseline, the bars would represent 68 and 74 percentage points. With a 60% baseline, their visible heights represent only 8 and 14 percentage points. The second visible bar is \(14/8=1.75\) times the first visible height, even though 74% is only about \(74/68\approx1.09\) times 68%. The truncated axis can therefore exaggerate the visual contrast. This is the kind of issue discussed in Misleading Axes and Scales in Bar Charts.

Critique the wording and conclusion. “Across the city” may imply that the display represents all city residents, but the chart gives results from 50 respondents in each neighborhood. Unless the survey design justifies that broader claim, the graph supports a description only of the respondents shown. Also, “overwhelmingly” is a vague characterization of the 74% value, and the survey percentages alone do not establish that the city should build the lane. That decision may require other information, and the graph does not show that support caused any outcome.

Write a more careful conclusion. Among the respondents represented in this survey, 74% of Westside respondents supported the protected bike lane, compared with 68% of Eastside respondents, a difference of 6 percentage points. The graph’s truncated vertical axis may make that difference appear larger than it is. This statement gives the observed values and limits the conclusion to the respondents represented.

A graph can contain useful data and still use a design that makes a comparison harder to judge. A good critique identifies the exact feature—here, the nonzero baseline—and explains its likely effect. Saying only “the graph is misleading” does not show that you understand why.

Common Mistakes and Full-Credit Communication

  • Naming a graph without explaining why it fits. State what variable or comparison the graph should display. For example, a bar chart suits a single categorical distribution; a segmented bar chart is useful for comparing conditional distributions across groups.
  • Reading a count graph as if it shows percentages. Check the axis label and the graph’s construction. Counts tell how many individuals are represented; percentages describe a share. For unequal group sizes, calculate within-group percentages before comparing how common a response is.
  • Assuming the tallest bar answers every question. A tallest bar identifies the largest displayed value, but the question might ask for the largest percentage, the smallest category, a group comparison, or a flaw. Match your answer to the wording.
  • Criticizing a graph without naming a specific flaw. Point to the feature: a missing title, unclear category labels, an unlabeled scale, a truncated bar-chart axis, or a display type that hides the needed comparison. Then say how that feature can confuse or mislead a reader.
  • Turning a sample description into an unsupported population claim. Use wording such as “among the surveyed participants” when that is all the graph establishes. Do not claim that a graph proves causation or that a difference is statistically significant.
  • Using vague evidence. Replace “the bars are very different” with the relevant values and a contextual comparison. If a value is estimated from the scale, use “about” or “approximately.”

A strong response to a mixed graph question is direct and specific. For a graph choice, name the display and say what comparison it makes clear. For a reading question, cite the category, group, value, and units or percentage basis. For a critique, identify a concrete issue and explain its effect. For a conclusion, stay within the data shown.

Key takeaway: Identify the question first, then check what the graph displays and whether that display fits the task. Support your answer with the right counts or percentages, and explain specific graph flaws without claiming more than the data show.

Check Your Understanding

For each question, explain your choice or conclusion and show any needed calculation.

  1. A survey of 120 visitors records their preferred exhibit among birds, fossils, insects, and plants. The museum wants to compare the numbers choosing each exhibit. Which graph would you recommend, and why?
  2. A chart shows 36 of 60 morning visitors and 45 of 90 afternoon visitors choosing a guided tour. Which group has the greater count, and which has the greater within-group percentage? Show both percentages.
  3. A bar chart comparing two categories starts its vertical scale at 40 rather than zero. Name the concern and explain how the scale may affect a reader.
  4. A segmented bar chart shows that 52% of one group and 47% of another group chose “often.” State the percentage-point difference and identify what denominator each percentage describes.
  5. A graph has clear bars but no title, and its horizontal-axis labels are abbreviated in a way that is not explained. Name two improvements and explain what each would clarify.