From Frequencies to Cumulative Percentages
In Choosing the Best Display for Quantitative Data, you compared graphs that show a distribution’s shape, individual values, or summary statistics. A cumulative relative frequency plot answers a different kind of question: what proportion of observations are below a given value, and what value marks a specified percentage of the data?
To build one, begin with a frequency table that groups quantitative data into consecutive intervals. Add the counts from the first interval through each interval in turn. These running totals are cumulative counts. Divide each cumulative count by the total number of observations to get cumulative relative frequencies. Plot those relative frequencies against the interval boundaries and connect the points with line segments.
The horizontal axis shows the quantitative variable in its original units. The vertical axis shows cumulative relative frequency, from 0 to 1, or cumulative percentage, from 0% to 100%. Because the values accumulate, the graph never decreases. Its points rise or stay level as you move from left to right, and the final cumulative relative frequency is 1, or 100%.
The earlier tutorials on relative frequencies and histograms introduced the ingredients for this graph. Here, the new step is to make the frequencies cumulative. A histogram shows how many observations are in each interval; a cumulative relative frequency plot shows how many have accumulated up to each successive boundary.
Construct the Plot Carefully
Suppose the intervals are half-open, such as \([0,10)\), \([10,20)\), and \([20,30)\). This notation includes the left endpoint and excludes the right endpoint. Thus, an observation equal to 10 belongs in \([10,20)\), not \([0,10)\).
For this convention, the cumulative count at the upper boundary of an interval includes observations in all intervals below that boundary. At 10, for example, it counts observations in \([0,10)\)—that is, values less than 10. It does not necessarily count values at or below 10, because observations exactly equal to 10 are assigned to the next interval. If there are observations exactly on a boundary, the plotted cumulative value is not the exact proportion at or below that boundary. Treat the graph as an estimate, especially when reading values between boundaries.
Worked Example: Build a Cumulative Relative Frequency Plot
Worked Example: Build a Cumulative Relative Frequency Plot
A fictional recreation center records the lengths of 50 visits, in minutes, and groups them into the intervals below. Construct the cumulative relative frequency table and identify the graph’s plotting points.
| Visit length (minutes) | Count | Cumulative count | Cumulative relative frequency |
|---|---|---|---|
| [0, 10) | 5 | 5 | 0.10 |
| [10, 20) | 10 | 15 | 0.30 |
| [20, 30) | 15 | 30 | 0.60 |
| [30, 40) | 12 | 42 | 0.84 |
| [40, 50) | 8 | 50 | 1.00 |
Plan. Add each interval’s count to the total from the intervals before it, then divide each running total by \(n=50\). Plot each result at its interval’s upper boundary.
Do. The cumulative counts are \(5\), \(5+10=15\), \(15+15=30\), \(30+12=42\), and \(42+8=50\). Dividing by 50 gives \(5/50=0.10\), \(15/50=0.30\), \(30/50=0.60\), \(42/50=0.84\), and \(50/50=1.00\). The plotted points at the upper boundaries are \((10,0.10)\), \((20,0.30)\), \((30,0.60)\), \((40,0.84)\), and \((50,1.00)\). Start the plot at \((0,0)\), assuming no observations fall below the first boundary of 0, and connect the points with line segments.
At 20 minutes, the plotted cumulative relative frequency is 0.30. With these half-open intervals, that represents the share of visits shorter than 20 minutes; visits exactly 20 minutes long are in \([20,30)\). The final point is 1.00 because all 50 observations fall below 50 minutes under the stated intervals.
Conclude. The graph rises from 0 to 1 as visit length increases. Its plotting points show the accumulated share at successive interval boundaries, using the interval convention to interpret each boundary correctly.
Read Percentiles and the Median
The \(p\)th percentile is a value at or below which about \(p\%\) of the observations fall. To estimate it from a cumulative relative frequency plot, locate \(p\%\) on the vertical axis, move horizontally to the graph, and then move vertically down to the horizontal axis. The corresponding horizontal value is the estimated percentile.
The median is the 50th percentile: it divides the ordered observations into two halves, with about half at or below it and about half at or above it. On a cumulative relative frequency plot, estimate the median by reading the horizontal value where the cumulative relative frequency is 0.50, or 50%.
When the target percentage falls between two plotted points, the graph uses a line segment to connect them. You can read an estimate from that segment, or calculate one by linear interpolation. Interpolation assumes the cumulative percentage increases evenly between the two boundaries. The data have been grouped, so the graph does not reveal the exact values within each interval. The result is an estimate, not a recovered observation.
Worked Example: Estimate a Median and a Percentile
Worked Example: Estimate a Median and a Percentile
A fictional clinic groups the time, in minutes, that 40 people spend in a waiting area. The cumulative relative frequencies at the upper boundaries are 0.10 at 5 minutes, 0.30 at 10 minutes, 0.60 at 15 minutes, 0.85 at 20 minutes, and 1.00 at 25 minutes. Estimate the median and the 75th percentile.
Plan. For each requested percentile, locate its cumulative relative frequency between the plotted points. Interpolate along the connecting line segment. Because the data are grouped, report approximate values.
Do: estimate the median. The median corresponds to 0.50, which lies between 0.30 at 10 minutes and 0.60 at 15 minutes. The increase from 0.30 to 0.50 is 0.20, out of a total increase of 0.60 to 0.30, or 0.30. That is \(0.20/0.30=2/3\) of the way from 10 to 15 minutes. The estimated median is therefore:
Do: estimate the 75th percentile. The cumulative relative frequency 0.75 falls between 0.60 at 15 minutes and 0.85 at 20 minutes. The fraction of the segment’s vertical rise is \((0.75-0.60)/(0.85-0.60)=0.15/0.25=0.60\). Thus, the estimate is \(15+0.60(20-15)=18\) minutes.
Conclude. The estimated median waiting-area time is about 13.3 minutes, and the estimated 75th percentile is about 18 minutes. These are estimates from grouped data and linear interpolation, not claims that an individual necessarily waited exactly either amount.
Estimate the Proportion Between Two Values
A cumulative relative frequency plot can also estimate the proportion of observations between two values. Read the cumulative relative frequency at each value, then subtract the lower cumulative estimate from the higher one. This works because the higher cumulative share includes observations below both cutoffs, while the lower share includes the observations below the first cutoff.
For cutoffs that fall inside intervals, this calculation inherits the graph’s interpolation assumption. It is therefore an estimate. Also, the exact inclusion of observations equal to either cutoff depends on how the data and intervals treat boundary values; grouped plots generally do not show enough detail to resolve ties exactly.
Worked Example: Estimate the Share in a Range
Worked Example: Estimate the Share in a Range
A fictional delivery service groups package drop-off times into intervals. The cumulative relative frequencies are 0.12 at 10 minutes, 0.40 at 20 minutes, 0.80 at 30 minutes, and 1.00 at 40 minutes. Estimate the percentage of drop-off times between 12 and 28 minutes.
Plan. Estimate the cumulative relative frequency at 12 minutes and at 28 minutes by linear interpolation between the neighboring plotted points. Subtract the first estimate from the second.
Do. At 12 minutes, the cumulative relative frequency is estimated between 0.12 at 10 minutes and 0.40 at 20 minutes. Twelve minutes is \(2/10\) of the way across that interval:
At 28 minutes, interpolate between 0.40 at 20 minutes and 0.80 at 30 minutes. Twenty-eight minutes is \(8/10\) of the way across:
Subtract the cumulative estimate at 12 from the one at 28: \(0.72-0.176=0.544\). As a percentage, \(0.544 \times 100\%=54.4\%\).
Conclude. The plot estimates that about 54.4% of the package drop-off times were between 12 and 28 minutes. Since both cutoffs are inside grouped intervals, interpolation makes this an estimate rather than an exact percentage from the individual data.
Common Mistakes and AP Exam Tips
- Using interval counts instead of cumulative counts. The graph’s vertical values must accumulate from the first interval onward. A single interval’s relative frequency is not the same as the cumulative relative frequency.
- Reversing the axes. Put the quantitative variable and its units on the horizontal axis. Put cumulative relative frequency or cumulative percentage on the vertical axis.
- Misreading a boundary as “at or below.” With half-open intervals such as \([10,20)\), the cumulative point at 20 represents the share below 20, not necessarily the share at or below 20. It equals the latter only if there are no observations exactly at 20. State what the plotted value represents under the interval convention.
- Forgetting that an interior reading is approximate. A line between grouped-data points does not show the original observations. Reading a percentile or cumulative share between boundaries assumes an even increase along that line.
- Subtracting in the wrong order for a range. For \(a<b\), subtract the cumulative relative frequency at \(a\) from the one at \(b\). The result estimates the share between the cutoffs.
- Calling the median an exact raw-data value. The 50% reading estimates the median from the grouped distribution. Unless the original observations are available, do not claim the exact middle observation is known.
For full-credit communication, name the quantity being estimated, include its units when it is a value of the quantitative variable, and identify the result as an estimate when it comes from interpolation. For a boundary reading, explain whether it represents values below the boundary or values at or below it under the interval convention.
Check Your Understanding
Use the graphing and reading ideas in this tutorial to answer each question.
- A grouped data set has 60 observations. The first three interval counts are 6, 15, and 18. What is the cumulative count and cumulative relative frequency through the third interval?
- For intervals \([0,5)\), \([5,10)\), and \([10,15)\), what does the cumulative relative frequency plotted at 10 represent? When would it also equal the proportion at or below 10?
- A cumulative relative frequency plot reaches 0.50 at about 24 units. What percentile does this estimate, and what familiar measure of center does it estimate?
- The cumulative relative frequency is estimated as 0.22 at 8 units and 0.67 at 18 units. Estimate the percentage of observations between 8 and 18 units.
- Why should an estimated 80th-percentile value read between two grouped-data boundaries be described as approximate?