Tutorials › AP Statistics › What to Do With an Unusual Point

Unusual points and model fit · Tutorial 971 of 1000

What to Do With an Unusual Point

Use a careful verification process to decide how an unusual point should be handled and communicate how it affects the regression.

Intermediate 10 min read

What You'll Learn

  • Separate an unusual point from a confirmed data-entry or measurement error.
  • Check original records, units, and case definitions before changing or excluding a value.
  • Compare regression summaries with and without a candidate point when that comparison is informative.
  • Report the effect on the fitted line, \(r^2\), and \(s\) in context.
  • Explain why a valid point should not be deleted simply because its removal improves fit.
  • Document corrections, exclusions, and sensitivity analyses transparently.

Unusual Does Not Mean Erroneous

In “Effect of Unusual Points on \(r^2\) and \(s\),” you saw that adding or removing a point can change both measures of fit. That comparison is useful, but it does not tell you whether the point is a mistake or whether it should be removed. A point can be unusual and still be a valid observation.

The first task is to investigate the case, not to make the regression look better. Check the original record, how the variables were measured, the units, and whether the case belongs in the data set. If the value is confirmed as an error, correct it using reliable information and explain the correction. If the value is valid, keep it in the primary analysis. You may also report a with-and-without comparison to show how much the result depends on it.

Definition: An unusual point is an observation that stands out in its position or effect relative to the other data. An unusual point is not automatically an error. A data error is a value that does not accurately represent the recorded case, such as a transcription mistake or a measurement recorded in the wrong units.

As in “Outliers, High-Leverage, and Influential Points Defined,” unusual position, a large residual, and influence are different descriptions. A point can have any one of these features without having the others. None of those labels, by itself, proves that the observation is wrong.

A Responsible Investigation

Start with the case itself. Confirm that the row is a real case and that the \(x\)- and \(y\)-values have been assigned to the correct variables. Check whether a decimal point, digit, unit conversion, or data-entry step could explain the unusual value. If available, compare the entry with the original form, instrument record, or other source. A value that looks implausible is a reason to check; it is not permission to replace it with a more convenient value.

Then decide how to handle what you learn. There are three common outcomes:

  • A specific error is confirmed and the correct value is known. Correct the value, document why, and analyze the corrected data. When the correction could materially affect the conclusion, show the result using the original entry as well.
  • The point is valid. Retain it in the primary analysis. If it is influential, a with-and-without comparison can show how the fitted model changes, but the comparison does not turn the point into an error.
  • The point cannot be verified or its status remains uncertain. Do not quietly alter or discard it. Describe the uncertainty and, when appropriate, report how the analysis changes with and without the point.

A with-and-without comparison is often called a sensitivity analysis: it checks whether the conclusions are sensitive to a particular observation or modeling choice. It is a way to show dependence on the point, not a rule for choosing the version with the most attractive \(r^2\), the smallest \(s\), or the most appealing slope.

Key practice: Verify the data first. Correct a confirmed error using evidence and document the correction. Retain valid observations in the primary analysis; if useful, report a clearly labeled with-and-without comparison as a sensitivity analysis.

What to Compare and Report

When you fit the regression with and without a candidate point, compare more than a single fit statistic. Look at the fitted line, including its slope and intercept, and at \(r\) or \(r^2\) and \(s\). As discussed in “Effect of Unusual Points on Slope” and “Effect of Unusual Points on Correlation,” the point can change the direction or strength of the association as well as the amount of scatter. The previous tutorial showed why \(r^2\) and \(s\) answer different questions: \(r^2\) is a proportion of response variation accounted for by the line, while \(s\) measures typical residual size in response units.

A clear report says which data are included in each fit and what changed. State the relevant results in context, with units for the slope, intercept, and \(s\). Do not describe a change in \(r^2\) as a change in the percentage of points on the line. And do not imply that a better-looking fit proves a point is invalid or establishes a causal relationship.

Worked Examples: Investigating and Reporting an Unusual Point

Worked Example: A Record Check Confirms a Typographical Error

Original AP-style question. In an invented transportation data set, \(x\) is a shuttle route’s distance in kilometers and \(y\) is its recorded travel time in minutes. Five rows are \((1,12)\), \((2,14)\), \((3,16)\), \((4,18)\), and \((5,80)\). The last travel time stands out. A check of the original dispatch record shows that it was 20 minutes, not 80. How should the result be handled and reported?

State. The value 80 appears unusual, but appearance alone would not justify changing it. Here, the original record confirms a transcription error and supplies the correct value, 20 minutes. We will report the fit using the entry as first recorded and the fit after the documented correction.

Plan. Compare the two regressions descriptively. There are no inference conditions to check for calculating these regression summaries. The key data-handling condition is that the correction is supported by the original record rather than chosen because it improves fit.

Do. With the original entry of 80, \(\bar{x}=3\) and \(\bar{y}=28\). The deviation sums are \(S_{xx}=10\), \(S_{yy}=3400\), and \(S_{xy}=140\). Thus the slope is \(140/10=14\), and the intercept is \(28-14(3)=-14\). The fitted line is \(\hat{y}=-14+14x\). Also,

$$ \mathrm{SSR}=\frac{140^2}{10}=1960, \qquad \mathrm{SSE}=3400-1960=1440 $$

With five observations, \(n-2=3\). Therefore \(r^2=1960/3400\approx0.5765\), and \(s=\sqrt{1440/3}=\sqrt{480}\approx21.9089\) minutes.

After the documented correction, all five points follow \(y=10+2x\). The fitted line is \(\hat{y}=10+2x\), with \(\mathrm{SSE}=0\), \(r^2=1\), and \(s=\sqrt{0/3}=0\) minutes. These exact-fit values result from the deliberately simple invented example.

Conclude. The record check—not the improvement in fit—justifies correcting 80 to 20. A transparent report would state that the recorded value was corrected from 80 to 20 minutes after checking the dispatch record, and would identify the corrected regression as the main analysis. If the initial entry is relevant to documenting the effect of the correction, the original-entry fit can also be reported. The \(r^2\) and \(s\) values describe these data; they do not show that route distance causes travel time.

Worked Example: A Valid Point Extends the Pattern

Original AP-style question. In an invented greenhouse exercise, \(x\) is hours of supplemental light per day and \(y\) is seedling growth in centimeters. Four observations are \((1,4)\), \((2,5)\), \((3,7)\), and \((4,8)\). A fifth seedling has \((8,14)\). The fifth case is confirmed as correctly measured. Compare the summaries, then decide whether the point should be removed.

State. The fifth observation has an \(x\)-value beyond the other four, so its position merits attention. It is confirmed as valid, however. We will use the first four observations as a descriptive comparison and retain all five in the primary analysis.

Plan. Calculate the regression summaries for the four-point fit, then refit with the confirmed fifth observation. Because this is a descriptive comparison, there are no inference conditions to check. The with-and-without comparison is a sensitivity analysis, not a test of whether the observation is erroneous.

Do. For the first four points, \(\bar{x}=2.5\), \(\bar{y}=6\), \(S_{xx}=5\), \(S_{yy}=10\), and \(S_{xy}=7\). The fitted line has slope \(7/5=1.4\) and intercept \(6-1.4(2.5)=2.5\). The fit summaries are

$$ \mathrm{SSR}=\frac{7^2}{5}=9.8, \qquad \mathrm{SSE}=10-9.8=0.2, \qquad r^2=\frac{9.8}{10}=0.9800, \qquad s=\sqrt{\frac{0.2}{4-2}}\approx0.3162\text{ cm} $$

For the added point, \(d_x=8-2.5=5.5\), \(d_y=14-6=8\), and the update factor is \(4/5=0.8\). Thus,

$$ S_{xx,\mathrm{new}}=5+0.8(5.5^2)=29.2, \qquad S_{yy,\mathrm{new}}=10+0.8(8^2)=61.2, \qquad S_{xy,\mathrm{new}}=7+0.8(5.5)(8)=42.2 $$

For all five observations, \(\mathrm{SSR}=42.2^2/29.2\approx60.9877\) and \(\mathrm{SSE}=61.2-60.9877\approx0.2123\). The slope is \(42.2/29.2\approx1.4452\); with \(\bar{x}=3.6\) and \(\bar{y}=7.6\), the intercept is approximately \(7.6-1.4452(3.6)=2.3973\). The fit measures are

$$ r^2=\frac{60.9877}{61.2}\approx0.9965, \qquad s=\sqrt{\frac{0.2123}{5-2}}\approx0.2660\text{ cm} $$

Conclude. Including the valid fifth seedling raises \(r^2\) from \(0.9800\) to about \(0.9965\) and lowers \(s\) from about \(0.3162\) cm to \(0.2660\) cm. The point extends the positive pattern and is close to the fitted line. It should remain in the primary analysis because it is a valid case, not because it improves the summaries. A report can give the all-five fit and, if useful, describe the four-point comparison as a sensitivity analysis.

Worked Example: A Valid Point Makes the Fit Less Strong

Original AP-style question. An invented delivery-planning data set records \(x\), distance from a distribution center in kilometers, and \(y\), delivery time in minutes. Four observations are \((2,8)\), \((4,11)\), \((6,10)\), and \((8,13)\). A fifth delivery, \((14,5)\), is unusual. The address and delivery time are verified. What should the analyst report?

State. The fifth point is far from the main pattern, but verification confirms it is a valid delivery. We will compare the fit without it to the fit with it, and retain it in the primary analysis unless there is a separate, stated reason that it falls outside the cases the model is meant to describe.

Plan. Calculate both fits and explain the changes in context. These summaries are descriptive, so there are no inference conditions to check. A fit that looks weaker after including the delivery is not evidence that its verified data should be deleted.

Do. For the first four observations, \(\bar{x}=5\), \(\bar{y}=10.5\), \(S_{xx}=20\), \(S_{yy}=13\), and \(S_{xy}=14\). The fitted line is \(\hat{y}=7+0.7x\). Since \(\mathrm{SSR}=14^2/20=9.8\), \(\mathrm{SSE}=13-9.8=3.2\), and there are two degrees of freedom,

$$ r^2=\frac{9.8}{13}\approx0.7538, \qquad s=\sqrt{\frac{3.2}{2}}\approx1.2649\text{ minutes} $$

For the fifth point, relative to the first four, \(d_x=14-5=9\), \(d_y=5-10.5=-5.5\), and the update factor is \(0.8\). The updated sums are

$$ S_{xx,\mathrm{new}}=20+0.8(9^2)=84.8, \qquad S_{yy,\mathrm{new}}=13+0.8((-5.5)^2)=37.2, \qquad S_{xy,\mathrm{new}}=14+0.8(9)(-5.5)=-25.6 $$

With all five observations, \(\mathrm{SSR}=(-25.6)^2/84.8\approx7.7283\) and \(\mathrm{SSE}=37.2-7.7283\approx29.4717\). The slope is \(-25.6/84.8\approx-0.3019\). Here \(\bar{x}=6.8\) and \(\bar{y}=9.4\), so the intercept is approximately \(9.4-(-0.3019)(6.8)=11.4528\). Therefore,

$$ r^2=\frac{7.7283}{37.2}\approx0.2078, \qquad s=\sqrt{\frac{29.4717}{5-2}}\approx3.1343\text{ minutes} $$

Conclude. Including the verified delivery changes the slope from \(0.7\) to about \(-0.3019\) minutes per kilometer, lowers \(r^2\) from about \(0.7538\) to \(0.2078\), and raises \(s\) from about \(1.2649\) to \(3.1343\) minutes. The association in the five-case data differs substantially from the four-case association. The appropriate report identifies this sensitivity and includes the valid delivery in the primary analysis; it does not discard the case merely to recover the stronger-looking four-point fit.

Common Mistakes and AP Exam Tip

  • Deleting a point because \(r^2\) rises or \(s\) falls. Those changes describe fit, not data validity. A full-credit explanation says whether the observation was checked and why it was corrected, retained, or excluded.
  • Calling every outlier a mistake. An unusual residual or \(x\)-position is a reason to investigate, not proof of a recording error. State what the data check established.
  • Changing a value without a source. Do not substitute a value that seems more reasonable or lies closer to the line. A correction needs evidence, such as a verified original record.
  • Reporting only the preferred fit. If the with-and-without comparison is important, identify both versions and explain the change. Do not conceal a sensitivity that affects the conclusion.
  • Confusing a sensitivity analysis with permission to exclude. A fit without a valid point can help show the point’s influence, but it is not automatically the correct primary analysis.
  • Leaving units or context out. Interpret a slope in response units per explanatory-variable unit and \(s\) in response units. Describe \(r^2\) as a proportion of response variation accounted for by the fitted line.

On an AP response, be explicit: identify the observation, explain what was checked, state whether the value was confirmed or corrected, and compare the fits if asked. If the point is valid, say that it remains in the primary analysis and describe the effect of including it. Keep the conclusion about model fit separate from any claim about cause.

Key takeaway: Investigate an unusual point before deciding how to handle it. Correct only a confirmed error, retain valid observations in the primary analysis, and use a clearly labeled with-and-without comparison to report sensitivity—not to choose whichever fit looks better.

Check Your Understanding

For each situation, focus on what the evidence supports and what a transparent report should say.

  1. A point has a large residual, but its original measurement record confirms the value. Is the point automatically an error? Explain.
  2. A source record confirms that a value was entered in the wrong units and provides the correct value. What should be documented when correcting it?
  3. Removing a valid observation raises \(r^2\) and lowers \(s\). What should happen to the primary analysis, and what comparison could still be reported?
  4. Why is a with-and-without comparison called a sensitivity analysis rather than proof that a point should be excluded?
  5. When reporting a comparison, what do \(r^2\) and \(s\) describe, and which one has response-variable units?