Tutorials › AP Statistics › Finding the Range of the Explanatory Variable

Extrapolation and prediction limits · Tutorial 942 of 1000

Finding the Range of the Explanatory Variable

Identify the actual minimum and maximum x-values used to fit a regression line so you can tell whether a requested x-value falls within or outside the data’s range.

Intermediate 9 min read

What You'll Learn

  • Identify which variable is the explanatory variable and locate its values in the data.
  • Find the minimum and maximum observed x-values, even when the data are unsorted.
  • State the observed x-range with the correct units and endpoints.
  • Check a requested x-value against the observed minimum and maximum.
  • Avoid using response values or a variable’s possible range in place of the observed x-range.
  • Confirm that the endpoints come from the same observations used to fit the regression model.

Why the Observed x-Range Matters

In “What Extrapolation Means,” we learned that whether a regression prediction is interpolation or extrapolation depends on the requested value of the explanatory variable. To make that comparison, you first need to find the smallest and largest explanatory-variable values among the observations used to fit the line.

That sounds straightforward, but it is easy to scan the wrong column, mistake the first and last entries for the endpoints, or use a range that describes what is possible rather than what was observed. This tutorial develops a reliable way to identify the observed range and report it clearly.

Definition: The observed range of the explanatory variable runs from its minimum observed value, \(x_{\min}\), to its maximum observed value, \(x_{\max}\), among the cases used to fit the regression model. In symbols, the endpoints are \(x_{\min}=\min(x_1,x_2,\ldots,x_n)\) and \(x_{\max}=\max(x_1,x_2,\ldots,x_n)\).

The notation \(x_1,x_2,\ldots,x_n\) represents the explanatory-variable values for the \(n\) cases in the data being used. The range is often reported as “from \(x_{\min}\) to \(x_{\max}\),” with units. It is the observed endpoints that matter, not whether the values between them were all observed.

Once you have the endpoints, compare the requested \(x\)-value with them. A requested value below \(x_{\min}\) or above \(x_{\max}\) is outside the observed range. A value between the endpoints, including an endpoint itself, is within the observed range. As the earlier tutorial explains, this comparison determines whether the prediction is extrapolation or interpolation.

A Reliable Method for Finding the Endpoints

Start by identifying the variables and their roles. In the regression context, \(x\) is the explanatory variable and \(y\) is the response variable. Use the explanatory variable’s measurements—not the response column—to find the range. For example, if a model predicts travel time from distance, distance is \(x\); the largest travel time does not tell you the largest observed distance.

Next, examine every \(x\)-value in the data used to fit the line. The entries may be unsorted, so do not assume the first value is the minimum or the last value is the maximum. Scan through the entire \(x\)-column, keeping track of the lowest and highest values. If the values are already in order from least to greatest, the first and last entries are the endpoints, but confirm that the order is actually ascending.

1
Locate \(x\).
Use the variable identified as explanatory in the regression setting, and include its units.
2
Check all relevant observations.
Scan the explanatory-variable values for the cases used to fit the line, even if they appear in no particular order.
3
Record both endpoints.
Write down the smallest observed value as \(x_{\min}\) and the largest as \(x_{\max}\).
4
Compare the requested value.
Decide whether it is below, within, or above the observed interval from \(x_{\min}\) to \(x_{\max}\).

A dataset may have repeated minimum or maximum values. That does not change the range: the endpoint is the value itself, whether it occurs once or several times. Nor does a gap in the observed values change the endpoints. The range summarizes the lowest and highest values; it does not say that every value between them appeared in the data.

Use the observations that were actually used to fit the model. If the regression analysis excludes cases—for example, because a measurement needed for the analysis is missing—do not automatically use those excluded cases when finding the model’s range. The range should describe the \(x\)-values represented in the fitted data. Likewise, a value that could have occurred in the setting but was not observed is not an endpoint of the observed range.

Worked Examples

Worked Example: Finding the Range in an Unsorted List

Fictional AP-style scenario. In a classroom activity, students record water temperature and dissolved oxygen in eight samples. Let \(x\) be water temperature in degrees Celsius. The temperatures, listed in the order the samples were measured, are \(18, 11, 24, 16, 26, 14, 21,\) and \(18\).

Identify the variable. The question is about the range of the explanatory variable, which is temperature \(x\), not dissolved oxygen \(y\). The temperature entries are not sorted, so the first value, 18, and last value, 18, cannot be assumed to be the endpoints.

Find the endpoints. Scanning all eight temperatures, the smallest is 11 and the largest is 26. The repeated value 18 does not affect either endpoint. Therefore, the observed \(x\)-range is from 11°C to 26°C.

Compare requested values. A request to predict at \(x=10\)°C is below the minimum of 11°C. It is outside the observed range. A request at \(x=26\)°C is exactly at the maximum endpoint, so it is within the range. A request at \(x=27\)°C is above the maximum and outside the range.

Conclude in context. The regression data include water temperatures from 11°C through 26°C. The requested temperature of 10°C is below the observed minimum, while 26°C is an observed endpoint and 27°C is above the observed maximum. This endpoint check is what allows each requested value to be classified.

Worked Example: Reading the Explanatory Column in a Table

Fictional AP-style scenario. A delivery service records distance and delivery time for eight fictional deliveries. Here \(x\) is distance in kilometers, and \(y\) is travel time in minutes. The table is organized by delivery, not by distance.

DeliveryDistance, \(x\) (km)Travel time, \(y\) (min)
A14,20038
B8,60025
C12,10034
D8,60028
E17,30049
F9,90030
G15,40044
H11,70032

Identify the variable. Distance is \(x\), the explanatory variable. The values in the travel-time column are responses, so the smallest and largest times are irrelevant to finding the \(x\)-range.

Find the endpoints. Scanning the distance column, the smallest value is 8,600 km and the largest is 17,300 km. The minimum occurs twice, but it remains a single endpoint value. The observed distance range is from 8,600 km to 17,300 km.

Compare requested values. A requested distance of 12,000 km is between the endpoints. A request of 17,300 km is at the maximum endpoint. A request of 18,000 km exceeds the maximum. Thus, the first two requested values are within the observed range, and the last is outside it.

Conclude in context. The fitted regression model is based on delivery distances from 8,600 km to 17,300 km. The range comes from the distance column, even though the table is not ordered by distance and the response values may look more naturally related to the question being asked.

Worked Example: Finding the Range When Values Are Already Ordered

Fictional AP-style scenario. A student records daily time spent using a language-learning app and a vocabulary quiz score for seven days. Let \(x\) be app use in hours per day. The \(x\)-values, already arranged in increasing order, are \(0.0, 1.0, 1.5, 2.0, 2.5, 3.5,\) and \(4.0\).

Find the endpoints. Since the values are explicitly in increasing order, the first entry is the minimum and the last entry is the maximum. Thus, \(x_{\min}=0.0\) hours per day and \(x_{\max}=4.0\) hours per day. The observed range is from 0 to 4 hours per day.

Check that the endpoints describe observations. A value of 0 is in the data: it means the student had a day with no app use. It is not merely a theoretical lower limit. Similarly, the largest observed value is 4 hours per day. If the app permits more than 4 hours of use, that possibility does not change the observed maximum.

Compare a request. A requested value of 3 hours per day lies between 0 and 4, even though 3 does not appear in the list. A request of 4.5 hours per day is above the maximum. The range is defined by the endpoints, not by whether a particular in-between value was recorded.

Conclude in context. For these seven days, the observed app-use values extend from 0 to 4 hours per day. A prediction at 3 hours per day is within that range; a prediction at 4.5 hours per day is outside it. The range says nothing by itself about how close a prediction will be to a quiz score.

What the Range Does—and Does Not—Tell You

The observed range answers a specific question: what are the smallest and largest explanatory-variable values represented in the data used to fit the line? It does not summarize the response values, tell you how many observations there are, or describe the strength of the relationship. Those are separate features of a regression analysis.

The endpoints also do not guarantee that a prediction within the range will be accurate. As discussed in “What Extrapolation Means,” being within the observed \(x\)-range makes a prediction interpolation; it does not ensure that the model describes the relationship well at every point. Conversely, identifying a value outside the range flags extrapolation, but does not by itself prove that a prediction is wrong.

Be precise about units. If \(x\) is measured in hours, report the minimum and maximum in hours. If a requested value is given in minutes, convert it to the same units before comparing. For instance, comparing 90 minutes directly with endpoints written in hours can lead to a mistaken classification. State the conversion or rewrite all values in one consistent unit.

Key takeaway: Find the observed \(x\)-range by locating the smallest and largest explanatory-variable values among the cases used to fit the model. Report both endpoints with units, then compare any requested \(x\)-value with that interval.

Common Mistakes and AP Exam Tips

  • Using the \(y\)-values instead of the \(x\)-values. The range for classifying a regression prediction is the range of the explanatory variable. A full-credit answer identifies which variable is \(x\) and gives its observed minimum and maximum.
  • Assuming the list is ordered. The first and last observations are not necessarily the minimum and maximum. Scan every \(x\)-value unless the data are clearly arranged from least to greatest or greatest to least.
  • Using a possible or intended range. A sensor’s operating limits, a survey’s allowed answers, or a variable’s physically possible values do not establish the observed range. Use the values in the data that were used to fit the regression model.
  • Ignoring the endpoints. The minimum and maximum themselves are within the observed range. A value is outside only if it is less than the minimum or greater than the maximum.
  • Forgetting units or mixing units. Report the range with the explanatory variable’s units, and convert a requested value before comparing if necessary.
  • Assuming every value between the endpoints was observed. The range gives the outer limits, not a complete list of the data. A value can be within the range even if it does not appear among the observations.
  • Using cases not included in the fitted data. When the problem specifies which observations were used to fit the model, base the endpoints on those observations. Explain the range as “the observed \(x\)-values used to fit the line,” rather than vaguely saying “the data range.”

A concise, complete response might read: “The explanatory variable is distance. The deliveries used to fit the model had distances from 8,600 km to 17,300 km, so those are the observed \(x\)-range endpoints. The requested distance of 18,000 km exceeds the maximum.” This names the variable, gives the endpoints and units, and makes the comparison explicit.

Check Your Understanding

For each question, identify the observed range of the explanatory variable and explain how you would compare a requested value with it.

  1. A regression uses weekly rainfall \(x\), in millimeters, with observed values \(12, 7, 19, 10,\) and \(15\). What are \(x_{\min}\) and \(x_{\max}\)?
  2. A table lists temperature as \(x\) and electricity use as \(y\). Which column should you scan to find the range used to classify a prediction for temperature?
  3. The observed explanatory-variable values run from 2.5 to 9.0 hours. Is a request at 9.0 hours within or outside the observed range? Explain.
  4. A data set has observed \(x\)-values of 4, 6, 6, 8, and 11. Does the repeated value 6 change the minimum or maximum? State the range.
  5. A variable could theoretically take values from 0 to 100, but the observations used to fit the model run only from 18 to 73. Which interval is the observed range, and why?