Tutorials › AP Statistics › Mixed Practice: Analyzing and Interpreting Data

Statistical practices and exam synthesis · Tutorial 1018 of 1020

Mixed Practice: Analyzing and Interpreting Data

Work through mixed data-analysis problems that combine graphs, numerical summaries, probability, and contextual interpretation.

Intermediate 10 min read

What You'll Learn

  • Describe a quantitative distribution using its shape, center, spread, and notable features.
  • Compare the mean and median to help interpret a distribution’s shape.
  • Calculate joint, inclusive-or, and conditional probabilities from a two-way table.
  • Read grouped data from a histogram and use counts to find empirical probabilities.
  • Explain what a calculated probability says in context without making unsupported claims.

Connect the Graph, the Calculation, and the Claim

Mixed data-analysis questions often move from one kind of reasoning to another. A prompt might ask you to describe a graph, calculate a probability from the data, and then explain what the result means. The key is to answer each part with the evidence it calls for. A graph describes a pattern; a calculation summarizes a feature or chance; an interpretation connects that result to the situation.

In “Matching Graphs and Summaries to Questions,” you practiced choosing evidence that fits the question. Here, you will apply that habit in a short mixed set. These problems focus on describing data and using observed counts to calculate probabilities—not on conducting an inference procedure. For each example, keep track of what the data represent and the group to which a probability applies.

Key takeaway: Move from evidence to calculation to interpretation. State what the graph or count shows, use the appropriate calculation, and explain the result in context without claiming more than the data support.

Mixed Practice Set

Try each prompt before reading its solution. For a useful practice session, set aside about 20 minutes: allow roughly 6 minutes for each prompt and a few minutes to review. As in “Approaching a Multi-Part Free-Response Question,” treat each requested part as a separate task and make each answer easy to find.

Worked Example: Describe a Distribution of Bus Wait Times

Prompt. A student records the wait, in minutes, for a bus on 12 different mornings. The ordered data are \(4, 5, 5, 6, 6, 7, 7, 8, 8, 9, 12, 18\). A histogram uses the intervals 0 to less than 5, 5 to less than 10, 10 to less than 15, and 15 to less than 20 minutes. Describe the distribution, calculate the mean and median, and explain which measure better represents a typical wait.

Step 1: Read the graph and describe the distribution. The histogram has 1 observation from 0 to less than 5 minutes, 9 from 5 to less than 10, 1 from 10 to less than 15, and 1 from 15 to less than 20. Most waits are between 5 and 10 minutes. The distribution has a concentration in that interval and a tail toward larger waits, including a wait of 18 minutes. This supports describing the distribution as right-skewed.

Wait interval (minutes)Frequency
0 to less than 51
5 to less than 109
10 to less than 151
15 to less than 201

Step 2: Calculate and compare center. The mean is the sum of the observations divided by the number of observations. The sum is \(95\), so the mean wait is:

$$ \bar{x}=\frac{95}{12}\approx 7.92\text{ minutes} $$

For 12 ordered observations, the median is the average of the sixth and seventh values. Both are 7, so the median wait is \((7+7)/2=7\) minutes. The mean is greater than the median, which is consistent with the right tail pulling the mean upward.

Step 3: Choose a representative measure and interpret. The median of 7 minutes is a reasonable measure of a typical wait because it is less affected by the unusually long 18-minute wait than the mean is. A complete interpretation is: “For these 12 recorded mornings, the median bus wait was 7 minutes.” The mean of about 7.92 minutes is also correct, but it summarizes these observations in a way that is more influenced by the longest wait.

Check the limits. The observations cover 12 mornings, not every possible bus trip or every rider’s wait. Describe the recorded data; do not claim that these values establish a typical wait for all routes, riders, or seasons.

Worked Example: Calculate Probabilities from a Two-Way Table

Prompt. In a hypothetical survey of 80 students, each student reports whether they usually arrive at school on time and how they usually travel. Of 50 students who usually take the bus, 36 arrive on time and 14 arrive late. Of 30 students who usually walk or bike, 27 arrive on time and 3 arrive late. Use the table to find the probability that a randomly selected surveyed student arrives late or takes the bus. Then find the probability that a student arrives late given that the student takes the bus, and interpret the comparison between travel groups.

Usual travel methodOn timeLateTotal
Bus361450
Walk or bike27330
Total631780

Step 1: Identify the events. Let \(L\) mean “the student arrives late” and \(B\) mean “the student takes the bus.” The table shows 17 students in \(L\), 50 in \(B\), and 14 in both \(L\) and \(B\).

Step 2: Calculate the probability of late or bus. The word “or” includes students who meet both conditions. Add the two event counts and subtract the overlap, which would otherwise be counted twice:

$$ P(L\text{ or }B)=\frac{17+50-14}{80}=\frac{53}{80}=0.6625 $$

Thus, the probability is 0.6625, or 66.25%. In context, among the 80 surveyed students, 66.25% either usually arrive late, usually take the bus, or do both.

Step 3: Calculate the conditional probability. “Given that the student takes the bus” changes the relevant group to the 50 bus riders. Of those, 14 arrive late:

$$ P(L\mid B)=\frac{14}{50}=0.28 $$

The conditional probability of arriving late among the surveyed bus riders is 0.28, or 28%. Among the 30 students who walk or bike, the late proportion is \(3/30=0.10\), or 10%. In this survey, late arrivals were more common among bus riders than among students who walked or biked. This describes an association in the surveyed students; it does not show that travel method caused the difference. Other factors, such as distance from school, could be related to both travel method and arrival time.

Worked Example: Combine a Histogram, Probability, and Interpretation

Prompt. A school records the wait, in minutes, for 40 riders who were randomly selected from its weekday shuttle riders. The histogram is summarized by the following class intervals and counts. Describe the distribution. Let \(X\) be the wait time for one rider randomly selected from these 40. Find \(P(X\geq 10)\) and \(P(5\leq X<15)\). Interpret both probabilities and state one limit on the conclusions.

Wait interval (minutes)FrequencyHistogram bar (each # represents 2 riders)
0 to less than 58####
5 to less than 1014#######
10 to less than 1510#####
15 to less than 206###
20 to less than 252#

Step 1: Describe the histogram. The frequencies rise from 8 to 14 and then decline to 10, 6, and 2. The highest bar is the 5-to-less-than-10-minute interval. The distribution has one main peak and a tail extending toward longer waits, so it is right-skewed. Most of the recorded waits are below 15 minutes, and only 2 of the 40 are at least 20 minutes.

Step 2: Define the random variable and calculate the first probability. \(X\) is the wait time, in minutes, for a rider selected at random from these 40 riders. The event \(X\geq 10\) includes the three intervals from 10 to less than 15, 15 to less than 20, and 20 to less than 25. Their counts total \(10+6+2=18\):

$$ P(X\geq 10)=\frac{18}{40}=0.45 $$

The probability is 0.45, or 45%. In context, if one of the 40 selected weekday riders is chosen at random, the probability that this rider’s recorded wait was at least 10 minutes is 0.45.

Step 3: Calculate the second probability. The event \(5\leq X<15\) includes the 5-to-less-than-10 and 10-to-less-than-15 intervals. Their counts total \(14+10=24\):

$$ P(5\leq X<15)=\frac{24}{40}=0.60 $$

The probability is 0.60, or 60%. In context, 60% of these 40 recorded waits were at least 5 minutes and less than 15 minutes.

Step 4: State a reasonable limit. The probabilities describe the 40 riders represented in this exercise. Random selection from weekday shuttle riders may support describing that group if the selection was carried out as planned, but the data do not automatically describe weekend riders or riders on other routes. They also do not explain why some waits were longer than others.

Check the Reasoning Before You Move On

A quick review can prevent a correct calculation from turning into an incomplete answer. Use the graph or table to identify the relevant evidence, check that the denominator matches the group in the question, and then write an interpretation that names that group. If the prompt asks for several things, make sure each one has its own response.

1
Read what is represented.
Identify the individuals, variables, units, and group shown in the graph or table.
2
Describe the pattern or identify the event.
For a distribution, consider shape, center, spread, and notable features. For a probability, translate the wording into the relevant count or counts.
3
Use the matching calculation.
For a proportion or empirical probability, divide the relevant count by the correct total. For “or,” account for any overlap.
4
Interpret and set a limit.
Report what the result means in context and avoid causal or population-wide claims that the data-collection plan does not support.

Common Mistakes and AP Exam Tips

  • Listing a shape without evidence. “The distribution is skewed” is stronger when you identify the tail’s direction and the data that support it. For example, point out that most waits are shorter while a few longer waits extend the right tail.
  • Choosing the mean automatically. The mean is useful, but a long tail or unusual value can pull it away from where most observations lie. Describe the distribution and compare the mean with the median before choosing a typical value.
  • Using the wrong denominator. For the probability of arriving late given that a student takes the bus, the denominator is the number of bus riders, not all surveyed students. Name the group after “given that” before calculating.
  • Counting overlap twice for “or.” In an inclusive “or” probability, students who meet both conditions belong to both event counts. Subtract the overlap once, or count directly how many students meet at least one condition.
  • Giving a number without its meaning. A probability of 0.45 is incomplete if the prompt asks for an interpretation. Say what event has probability 0.45 and identify the relevant group, such as the selected weekday shuttle riders.
  • Turning an observed difference into a cause. A larger late-arrival proportion among bus riders describes an association in the survey. It does not establish that taking the bus caused lateness.
  • Claiming more than the data cover. If the data concern selected weekday riders on one shuttle, do not silently extend the conclusion to every student or every route. Match the claim to the people and conditions represented.

For full credit, show enough work to make the reasoning visible: identify the group or event, show the counts used, give the calculated result with suitable rounding, and explain it in context. A careful sentence about what the data show is more valuable than an extra claim the data cannot support.

Key takeaway: In mixed data analysis, keep the graph, calculation, and interpretation connected. Use the right group and denominator, describe patterns with evidence, and keep conclusions within the scope of the data.

Check Your Understanding

Try these questions without looking back at the worked examples. For each probability, state what group the denominator represents.

  1. A set of 10 ordered quiz scores is \(3, 5, 6, 6, 7, 7, 8, 8, 9, 11\). Find the median and describe the distribution’s shape using the data.
  2. In a group of 60 students, 22 take the late bus, 18 attend an after-school club, and 8 do both. Find the probability that a randomly selected student takes the late bus or attends a club.
  3. In the same group, 8 of the 22 late-bus riders attend a club. Find the probability that a student attends a club given that the student takes the late bus.
  4. A histogram shows 12 observations from 0 to less than 10 minutes and 8 observations from 10 to less than 20 minutes. What is the empirical probability that a randomly selected observation is at least 10 minutes?
  5. A survey finds different proportions of late arrivals for two travel methods. State one conclusion the comparison supports and one conclusion it does not establish by itself.