From a List of Observations to a Useful Summary
A statistics question may give you a list of raw data rather than a graph or a table of summary values. You need to organize the observations, calculate the requested measures, and explain what the results mean for the group and variable. A correct number without a clear interpretation is only part of a complete response.
The tutorials on computing the mean, finding the median and quartiles, calculating standard deviation, and interpreting percentiles and z-scores established the individual tools. Here, you will bring them together in a full exam-style problem. The main new habit is to use a consistent workflow: check the data, order them, calculate position-based summaries, calculate mean-based summaries, and then interpret each result with units.
A Reliable Workflow for Raw Data
Start by checking that you have included every observation and identified the variable and its units. Sort the values from least to greatest before finding the median, quartiles, or percentile rank. Keep the original values available when calculating the mean and standard deviation.
Name the group, the quantitative variable, its units, and the number of observations \(n\). Check that the list contains \(n\) values.
Sort the values from least to greatest. Use this list to locate the median, quartiles, and the number of observations at or below a specified value.
Find the mean and median, then calculate requested measures such as range, IQR, and sample standard deviation \(s\). Keep units attached to the results.
Use a percentile rank to report a share of observations at or below a value. Use a z-score to report how many standard deviations the value is from the mean.
For quartiles, follow the convention used in this course and in Finding Quartiles and the IQR: split the ordered values into lower and upper halves, then find the median of each half. For an even number of observations, the halves have the same size. For an odd number, leave the overall median out of both halves.
Worked Example: A Full Summary of Reading Time
Worked Example: A Full Summary of Reading Time
A fictional group of 12 students recorded how many minutes they spent reading on one evening. The observations, in minutes, were \(14,16,18,19,20,21,22,23,24,25,27,31\). Calculate and interpret the mean, median, quartiles, range, IQR, sample standard deviation, the percentile rank of 24 minutes, and the z-score of 31 minutes.
State. The group is these 12 students, the variable is evening reading time, and the units are minutes. We will summarize the observed times; these results describe this group, not all students.
Plan. The observations are already ordered, and there are 12 of them. Split the list into two halves of six to find \(Q_1\) and \(Q_3\). Use the sample standard deviation because these observations are a sample of students, not every student in a defined population. Use the at-or-below convention for percentile rank, as in Percentiles and Their Interpretation.
Do: center. Add the observations to get \(260\) minutes. Divide by the sample size for the mean:
A check is that \(12(21.6667)\) is approximately 260, the total of the observations. With an even sample size, the median is the average of the sixth and seventh ordered values, 21 and 22:
Do: position-based spread. The lower half is \(14,16,18,19,20,21\), so \(Q_1=(18+19)/2=18.5\) minutes. The upper half is \(22,23,24,25,27,31\), so \(Q_3=(24+25)/2=24.5\) minutes. Therefore:
The range is the maximum minus the minimum: \(31-14=17\) minutes. The IQR describes the width of the middle half of the observations, from 18.5 to 24.5 minutes; the range describes the full span, from 14 to 31 minutes.
Do: sample standard deviation. The sample standard deviation uses the squared deviations from the mean and divides their sum by \(n-1\) before taking the square root. Using \(\sum x=260\) and \(\sum x^2=5882\), the squared-deviation sum is \(5882-260^2/12=746/3\). Thus:
The calculation can be checked by noting that the sample variance is \(746/33\approx22.6061\), and its square root is about \(4.7546\). In context, the reading times in this group typically differed from their mean of about 21.67 minutes by around 4.75 minutes.
Do: position. Nine of the 12 observations are 24 minutes or less. The percentile rank of 24 minutes under the at-or-below convention is:
The z-score of 31 minutes describes its distance from the mean in sample-standard-deviation units:
The positive z-score means 31 minutes is about 1.96 sample standard deviations above the group’s mean. It is a standardized distance, not a percentile rank.
Conclude in context. For these 12 students, mean evening reading time was about 21.67 minutes and median time was 21.5 minutes. The middle half of their times extended from 18.5 to 24.5 minutes, a width of 6 minutes; the full range was 17 minutes. The sample standard deviation was about 4.75 minutes. A time of 24 minutes was at or above 75% of the observed times, while 31 minutes was about 1.96 standard deviations above the mean.
Worked Example: When Center Measures Differ
Worked Example: Time to Finish a Puzzle
In a fictional group of eight students, the times to finish a puzzle were \(6,7,8,8,9,10,10,14\) minutes. Find the mean, median, quartiles, IQR, range, and sample standard deviation. Then give a suitable interpretation of center and spread.
The sum is \(72\), so the mean is \(72/8=9\) minutes. The median is the average of the fourth and fifth observations: \((8+9)/2=8.5\) minutes. The lower half is \(6,7,8,8\), giving \(Q_1=(7+8)/2=7.5\) minutes. The upper half is \(9,10,10,14\), giving \(Q_3=(10+10)/2=10\) minutes. Therefore, the IQR is \(10-7.5=2.5\) minutes. The range is \(14-6=8\) minutes.
For the sample standard deviation, the deviations from the mean of 9 are \(-3,-2,-1,-1,0,1,1,5\). Their squares sum to \(42\), so:
A check is that the sample variance is \(42/7=6\), whose square root is about 2.45. The mean of 9 minutes is above the median of 8.5 minutes, and the largest time is 14 minutes. As discussed in Choosing Mean or Median to Describe Center, the median and IQR are useful resistant summaries when a high value may pull the mean or affect the standard deviation. In context, the median puzzle time was 8.5 minutes, and the middle half of times extended from 7.5 to 10 minutes, a width of 2.5 minutes.
Worked Example: Percentile Rank and Standardized Position
Worked Example: Comparing a Trail Run Time with the Group
Ten fictional runners recorded times, in minutes, on the same short trail: \(18,19,20,20,21,22,23,24,26,27\). Find the mean and sample standard deviation, then describe the percentile rank of 24 minutes and the z-score of 27 minutes.
The sum is \(220\), so the mean is \(220/10=22\) minutes. To calculate the sample standard deviation, the squared deviations from 22 are \(16,9,4,4,1,0,1,4,16,25\), with sum \(80\). Therefore:
Eight of the ten times are 24 minutes or less, so the percentile rank of 24 minutes is \((8/10)\times100\%=80\%\). Under the at-or-below convention, about 80% of these observed times were 24 minutes or less. This is a statement about the group’s times, not the percentage of a race completed.
For 27 minutes, the z-score is:
Thus, 27 minutes is about 1.68 sample standard deviations above the mean time. Since a longer time means a slower run, this runner’s time was above the group’s mean time. The z-score does not automatically give a percentile; it expresses distance from the mean in standard-deviation units.
Common Mistakes and AP Exam Tips
- Calculating position summaries before ordering the data. An unsorted list can lead to the wrong median, quartiles, or count at or below a value. Show or create the ordered list first.
- Mixing up \(Q_1\), \(Q_3\), and the IQR. Quartiles are values in the data’s scale; the IQR is their difference. State both quartile endpoints when interpreting the middle half, then report the width in the original units.
- Using \(n\) instead of \(n-1\) for a sample standard deviation. For sample data, calculate \(s\) with \(n-1\) in the denominator. Use \(Sx\), not \(\sigma x\), when reading sample standard deviation from calculator 1-Var Stats output.
- Describing standard deviation as the full spread. Standard deviation is a typical distance from the mean, not the distance between the minimum and maximum. Use “range” for the full span.
- Calling a percentile rank a z-score, or vice versa. A percentile rank is a percentage of observations at or below a value. A z-score is a signed distance from the mean in standard-deviation units.
- Reporting numbers without context. “The IQR is 6” is incomplete. A full-credit statement identifies the group, variable, value, and units: “The middle half of these students’ reading times spans 6 minutes.”
- Reporting too many summaries without choosing a useful interpretation. As emphasized in Choosing Appropriate Summary Statistics, inspect the distribution and any unusual values. For skewed data or outliers, the median and IQR are often more useful; for approximately symmetric data without strong outliers, the mean and standard deviation are often appropriate.
For a full-credit exam response, make the requested calculations visible and label the statistics. Then write an interpretation that answers the question in context. Do not imply that a sample summary automatically describes a larger population, and do not convert a z-score into a percentile unless the problem provides an appropriate model and asks for that calculation.
Check Your Understanding
Use the data and contexts given in each question. Show calculations and state interpretations in context.
- A fictional group of seven students recorded the number of minutes spent practicing an instrument: \(12,15,15,18,20,22,26\). Find the mean, median, range, and sample standard deviation.
- For the ordered data \(4,6,7,9,10,12,14,17\), find \(Q_1\), \(Q_3\), and the IQR using the median-of-halves convention. Interpret the IQR.
- In a group of 20 measurements, 15 are less than or equal to 8. What is the percentile rank of 8 under the at-or-below convention? Write what it means.
- A measurement is 3 units below a group mean, and the group’s standard deviation is 2 units. Find its z-score and interpret its sign and magnitude.
- Explain why a percentile rank and a z-score are different descriptions of a value’s position, even when they refer to the same data set.