An Outlier Is a Clue, Not a Verdict
A graph or the 1.5 IQR rule can identify an observation that stands apart from the rest. But the display alone cannot explain why that value is unusual. It might be a recording or measurement error, it might come from a different group or process, or it might be an accurate observation of a genuinely rare event.
In Identifying Outliers and Unusual Features, you learned to describe what a graph shows and to use the 1.5 IQR rule to flag possible outliers. This tutorial builds on that distinction: a statistical flag describes a value’s position relative to the other observations. It does not, by itself, tell you what caused the value or what to do with it.
The context supplies questions that the numerical rule cannot answer. Was the measurement made and recorded correctly? Did the observation come from the same population and process as the others? Is there evidence that an unusual but real event occurred? Until those questions have been investigated, treat the observation as a value to check—not as a value to erase.
Three Possible Explanations
An unusually high or low value can have different explanations. These possibilities call for different responses, so do not settle on one based only on how far the value lies from the rest of the data.
- Error: The value may have been entered incorrectly, measured with a faulty instrument, recorded in the wrong units, or attached to the wrong observation. An unusual value is a reason to check the source, not evidence by itself that an error occurred.
- A different group or process: The observation may come from a different population, setting, machine, method, or time period. Combining distinct groups can make one group’s ordinary values look unusual relative to the combined data. Investigate whether the distinction is supported by information about how the observations were collected.
- A true extreme: The value may be accurate and come from the same relevant process, while still being rare or extreme. Such a value can be important. For example, an unusually large environmental measurement might reflect a real event that the data are intended to capture.
These explanations are not always immediately distinguishable. A single unusual measurement might lead you to ask about the instrument, the observation’s group, and whether a rare event occurred. The evidence—not the outlier rule alone—should guide your conclusion.
A Responsible Investigation
When a value stands apart, describe what you can see, then investigate what the data collection context can establish. The following sequence helps keep those tasks separate. It is a reasoning routine, not a new statistical test.
State which observation is unusually high or low and, if relevant, whether it falls outside a fence. Say “flagged as a possible outlier” when the cause is not known.
Review the original record, units, measurement procedure, instrument, and data-entry steps. Seek independent confirmation when available. Do not assume an error just because a value is surprising.
Ask whether it was collected from the same kind of subject, group, location, equipment, time period, and procedure. Use documented information, not a convenient explanation invented after seeing the value.
Correct a verified recording error according to the data-handling rules and document the correction. If groups differ meaningfully, describe or analyze the groups in a way that fits the question. If the value is genuine and relevant, retain it and explain what it represents.
The 1.5 IQR rule remains useful here, but its role is limited. As established in Applying the 1.5 IQR Rule for Outliers, calculate the fences from \(Q_1\), \(Q_3\), and the IQR; observations strictly outside a fence are flagged. A fence is a statistical reference point, not a boundary between “impossible” and “possible” values. A genuine observation can fall outside it, and an error can fall inside it.
The appropriate response also depends on the question. If the goal is to describe all observed events, a genuine extreme may be central to the story. If the goal is to compare a particular group, observations from a different documented group may need to be described separately. If a value is confirmed to be a recording error, it is not an accurate observation of the variable as intended. In each case, explain the evidence behind the decision.
Worked Examples
Worked Example: Checking a Surprisingly Long Service Time
A fictional repair desk records service times, in minutes, of 11, 12, 12, 13, 14, 14, 15, 16, and 61. The 61-minute time appears far above the others. Use the 1.5 IQR rule to flag it, then interpret what the flag does and does not establish.
State. The variable is the service time for a repair-desk visit, measured in minutes. The observation of 61 minutes is much larger than the other eight values, but we need to check the quartile fences and the record before explaining why.
Plan. The values are in order, and there are \(n=9\) observations. Use the median-of-halves convention from the earlier work on the five-number summary: find \(Q_1\) and \(Q_3\) from the lower and upper halves, calculate the IQR and fences, and then investigate the unusual record.
Do. The overall median is 14. The lower half is 11, 12, 12, 13, so \(Q_1=(12+12)/2=12\). The upper half is 14, 15, 16, 61, so \(Q_3=(15+16)/2=15.5\). Thus the IQR is \(15.5-12=3.5\) minutes. The lower fence is \(12-1.5(3.5)=6.75\) minutes, and the upper fence is \(15.5+1.5(3.5)=20.75\) minutes. Since 61 is greater than 20.75, it is flagged as a possible high outlier.
A check of the fictional desk’s original service log shows that the visit took 16 minutes; 61 was a transposition error when the time was entered. The evidence supports treating this as a recording error, not as an unusually long visit. A responsible data handler would follow the established correction procedure and keep a record of what was changed and why.
Conclude. The value 61 minutes was outside the upper fence, so it merited investigation. The original log confirms it was recorded incorrectly: the verified service time was 16 minutes. The fence alone did not reveal that cause; checking the source did.
Worked Example: A Value From Different Equipment
A fictional lab combines processing times, in minutes, from ten runs: 10, 11, 12, 12, 13, 14, 14, 15, 16, and 37. The log notes that the 37-minute run used a different machine from the other nine. Decide what can be said about the flag and what further information is needed.
State. The quantitative variable is processing time in minutes. The 37-minute observation is unusually high, and the equipment note gives a possible reason to check whether the runs are comparable.
Plan. Apply the 1.5 IQR rule to the combined list to confirm the statistical flag. Then use the machine information carefully: a different machine is a documented distinction, but one observation alone does not establish how that machine generally performs.
Do. There are ten ordered values. The lower half is 10, 11, 12, 12, 13, so \(Q_1=12\). The upper half is 14, 14, 15, 16, 37, so \(Q_3=15\). Therefore, \(\mathrm{IQR}=15-12=3\) minutes. The upper fence is \(15+1.5(3)=19.5\) minutes; the lower fence is \(12-1.5(3)=7.5\) minutes. Because \(37>19.5\), the run is flagged as a possible high outlier in the combined distribution.
The equipment note makes machine type a relevant variable to investigate. Check whether the machine label is correct, whether the same timing procedure was used, and whether there are other runs on each machine. If the machines represent distinct processes and the question concerns their performance, compare or describe those groups separately using adequate data. With only one run on the second machine, do not claim that its typical processing time is 37 minutes or that the machine caused the long time.
Conclude. The 37-minute run is outside the combined data’s upper fence. It is also documented as coming from different equipment, so the pooled data may combine observations from distinct processes. The machine record motivates a subgroup check, but additional comparable runs and a consistent procedure are needed to characterize a machine’s processing times.
Worked Example: Retaining a Verified Extreme Event
A fictional monitoring station records water levels, in centimeters, on 11 observation days: 11, 12, 12, 13, 13, 14, 14, 15, 16, 17, and 45. A storm occurred on the day with the 45-centimeter reading, and a second calibrated gauge confirms the measurement. Explain how to interpret the high value.
State. The variable is the water level recorded at the station, in centimeters. The observation of 45 centimeters is much higher than the other readings. In this scenario, the storm record and second gauge provide evidence that it is a genuine measurement.
Plan. Use the 1.5 IQR rule to identify whether 45 is flagged, then interpret the available context. Because the question concerns the observed water levels, a verified storm-related reading may be important rather than an error to discard.
Do. There are 11 values, and the median is the sixth value, 14. The lower half is 11, 12, 12, 13, 13, so \(Q_1=12\). The upper half is 14, 15, 16, 17, 45, so \(Q_3=16\). The IQR is \(16-12=4\) centimeters. The upper fence is \(16+1.5(4)=22\) centimeters, and \(45>22\), so 45 is flagged as a possible high outlier.
The outlier flag identifies an observation that stands apart from the other recorded levels. It does not contradict the independent gauge or the storm information. Those facts support interpreting the value as a real, unusually high water level associated with the storm. If the purpose is to describe levels observed on all days, omitting it would conceal a relevant event.
Conclude. The 45-centimeter reading is outside the upper fence, but the stated evidence supports treating it as a genuine extreme observation, not a recording error. Describe it in context as a high water level recorded during the storm, and make clear that it is unusual relative to the other observations.
Common Mistakes and AP Exam Tips
- Calling every flagged value an error. A full-credit response distinguishes the statistical flag from its cause: “The value is outside the upper fence and is flagged as a possible outlier; the rule alone does not show whether it is an error.”
- Removing a value simply because it is inconvenient. An extreme value can be accurate and relevant. Before changing or excluding an observation, identify evidence and explain why the decision fits the data and the question.
- Inventing a subgroup explanation. A value might come from a different group, but that claim needs contextual evidence such as a documented machine, population, location, or collection method. A plausible story is not proof.
- Treating a corrected record as if no decision occurred. When a source confirms a recording error, follow the stated data-handling procedure and document the correction. Keep the distinction clear between the recorded value and the verified value.
- Reporting a fence without units or context. Fences have the same units as the quantitative variable. Name the variable and use the fence to support a statement about the particular observation.
- Confusing unusual with impossible. A fence is a rule for flagging values relative to a distribution, not a physical limit. Check whether the value is plausible and supported by the collection process.
Check Your Understanding
For each situation, separate what the statistical flag establishes from what the context might establish.
- A sample has \(Q_1=8\), \(Q_3=12\), and an observation of 20. Find the IQR and upper fence. What can you conclude if 20 is above that fence?
- A scale reading is much higher than the rest, but no one has checked the scale or the original record. What is a careful way to describe the value, and what should be checked?
- A flagged observation is labeled as coming from a different location. What additional question should be considered before claiming that location explains the value?
- A very high rainfall measurement is confirmed by a second instrument and coincides with a documented storm. Explain why the value might be important to retain.
- Why is “This point is an outlier, so it must be wrong” not a justified conclusion from the 1.5 IQR rule?