Tutorials › AP Statistics › Predictions From Computer Output

Least-squares regression · Tutorial 876 of 1000

Predictions From Computer Output

Read the coefficient estimates in computer output, form the correct regression equation, and use it to predict a response for a specified predictor value.

Intermediate 9 min read

What You'll Learn

  • Identify the intercept and predictor coefficient in a computer-output table.
  • Distinguish coefficient estimates from other columns in the output.
  • Translate software labels into a regression equation.
  • Substitute a requested predictor value and calculate the predicted response.
  • Check that the equation, predictor, and response match the question.
  • Explain what a computed value represents without treating it as an observed outcome.

From Computer Output to a Prediction

In “Making Predictions Using the Equation,” you practiced substituting a value into a regression equation. Computer output often gives you the coefficients without displaying the equation in the form you have practiced. The essential task is to find the estimate for the intercept and the estimate for the predictor, put them in the correct places, and evaluate the equation at the requested predictor value.

A coefficient table can contain several columns and rows. The row labeled “(Intercept)” gives the intercept \(a\). The row for the predictor gives the slope \(b\), and the “Estimate” column gives the coefficient values to use. Other columns—such as standard errors, test statistics, or p-values—are not coefficients to substitute when making a prediction.

Definition: For a simple linear regression output table, the estimated intercept is the coefficient estimate in the “(Intercept)” row, and the estimated slope is the coefficient estimate in the predictor’s row. Use them to write \(\hat{y}=a+bx\), where \(x\) is the predictor and \(\hat{y}\) is the predicted response.

The names shown in output can differ from the variable names you use in an explanation. For example, software might display “distance_km” as a row label. The context tells you whether that variable is the predictor or the response. As covered in “Reading Regression Computer Output” and “Writing the Equation With Variable Names,” check the row label and the coefficient column before writing an equation.

Once you have the equation, use the predictor value specified in the question. Substitute that value for \(x\), multiply the slope by it, and then add the intercept. The result is a predicted response, \(\hat{y}\), not necessarily the response that any particular individual will actually have.

Worked Example: Predicting a Commute Time

Suppose a fictional transportation class collects data on the distance of a commute and the time it takes. Let \(x\) be commute distance in kilometers and \(y\) be commute time in minutes. The following invented computer output reports the coefficient estimates from a simple linear regression of time on distance.

TermEstimate
(Intercept)6.35
distance_km2.18

Worked Example: Predicting a Commute Time

Use the output to predict the commute time for a distance of 8.5 kilometers.

State. The predictor is commute distance, \(x\), measured in kilometers. The response is commute time, \(y\), measured in minutes. We need the regression line’s predicted time when \(x=8.5\).

Plan. Take the intercept from the “(Intercept)” row and the slope from the “distance_km” row, both in the “Estimate” column. Write \(\hat{y}=a+bx\), then substitute 8.5 for \(x\).

Do. The intercept is 6.35, and the slope is 2.18. Therefore, the fitted line is:

$$ \widehat{\text{time}}=6.35+2.18(\text{distance}). $$

Substitute a distance of 8.5 kilometers:

$$ \hat{y}=6.35+2.18(8.5)=6.35+18.53=24.88\text{ minutes}. $$

Conclude. For a commute distance of 8.5 kilometers, the regression line predicts a commute time of 24.88 minutes. This is the line’s predicted time; an individual commute could take more or less time.

A useful check is to confirm that the coefficient came from the right row and that the requested value was substituted for the predictor. Here, the input is distance in kilometers, not time in minutes. The calculation also has the form intercept plus slope times predictor, matching the equation \(\hat{y}=a+bx\).

Read the Row and Column Before Calculating

The coefficient table may also contain columns such as “Std. Error,” “t value,” or “Pr(>|t|).” These columns have other purposes; they are not the estimated intercept or slope. For a prediction, locate the “Estimate” column and then match each estimate to its row. Do not choose a number merely because it is next to the predictor name.

The response variable may be identified in a heading, a model formula, or the surrounding description rather than in a coefficient row. A simple regression coefficient table normally lists the intercept and predictor terms; it does not list the response as a coefficient. Before calculating, use the situation or output heading to establish which variable is predicted and which value is being supplied as the predictor.

Worked Example: Predicting Basil Growth

A fictional greenhouse class models the dry mass of basil plants using the number of hours of light per day. Let \(x\) be daily light hours and \(y\) be dry mass in grams. The invented output is:

TermEstimateStd. Error
(Intercept)12.43.1
light_hours3.60.5

Worked Example: Predicting Basil Growth

Use the output to predict the dry mass of a plant receiving 7.5 hours of light per day.

State. Daily light hours is the predictor, and plant dry mass is the response. The requested input is \(x=7.5\) hours per day.

Plan. Use the estimates, not the standard errors. The intercept estimate is 12.4 from the “(Intercept)” row, and the slope estimate is 3.6 from the “light_hours” row. Substitute the light value into \(\hat{y}=a+bx\).

Do. The fitted equation is:

$$ \widehat{\text{dry mass}}=12.4+3.6(\text{light hours}). $$

At 7.5 hours per day, the prediction is:

$$ \hat{y}=12.4+3.6(7.5)=12.4+27=39.4\text{ grams}. $$

Conclude. The model predicts a dry mass of 39.4 grams for a basil plant receiving 7.5 hours of light per day. The standard errors, 3.1 and 0.5, are not substituted into the prediction equation; they are not coefficient estimates.

This example highlights a common output-reading trap. The 0.5 in the predictor row is a standard error, not the slope. Substituting it would produce a different equation and an incorrect prediction. When the table has several numerical columns, name the column you are reading before you write down a coefficient.

Worked Example: Predicting a Swim Time

In another invented example, a coach records weekly training minutes and 100-meter swim times for a group of swimmers. Let \(x\) be weekly training minutes and \(y\) be swim time in seconds. The output gives these coefficient estimates:

TermEstimate
(Intercept)94.60
weekly_training_min-0.31

Worked Example: Predicting a Swim Time

Find the predicted 100-meter swim time for a swimmer with 40 minutes of weekly training.

State. Weekly training time is the predictor, and 100-meter swim time is the response. We want the predicted response at \(x=40\) minutes.

Plan. Read the intercept estimate as 94.60 and the predictor estimate as \(-0.31\). Keep the negative sign when writing the line, then substitute 40 for weekly training minutes.

Do. The fitted equation is:

$$ \widehat{\text{swim time}}=94.60-0.31(\text{weekly training minutes}). $$

Evaluate the line at 40 minutes:

$$ \hat{y}=94.60+(-0.31)(40)=94.60-12.40=82.20\text{ seconds}. $$

Conclude. For a swimmer with 40 minutes of weekly training, the line predicts a 100-meter time of 82.20 seconds. This is a prediction from an observed association, not evidence that increasing training by a particular amount will cause an individual swimmer’s time to change by that amount.

A negative estimate is still used in the same equation, \(\hat{y}=a+bx\). You can write it with a plus sign followed by a negative slope, or with a minus sign and the slope’s positive magnitude. Either way, retain the sign in the output. This agrees with the direction of the fitted line described in “Positive and Negative Slopes in Context.”

A Reliable Sequence for Computer-Output Predictions

A short, repeatable sequence helps prevent mistakes when the output is unfamiliar. Identify the response and predictor from the context. Find the “Estimate” column, then take the intercept from its row and the slope from the predictor’s row. Write the equation before substituting, so you can see which coefficient multiplies the requested value. Finally, evaluate the expression and describe the result as a predicted response in context.

1
Identify the variables.
Establish which variable is the predictor \(x\), which is the response \(y\), and what value of \(x\) the question supplies.
2
Locate the coefficient estimates.
Read the intercept from the “(Intercept)” row and the slope from the predictor’s row in the “Estimate” column.
3
Write and evaluate the line.
Form \(\hat{y}=a+bx\), substitute the requested predictor value, and calculate the predicted response.
4
State the prediction in context.
Name the response and the predictor value. Make clear that the answer is predicted, not guaranteed to be an observed outcome.

Common Mistakes and AP Exam Tips

  • Using a standard error instead of an estimate. In a table with multiple columns, the estimate column supplies the coefficients. A standard error describes uncertainty about an estimate; it does not replace that coefficient in the prediction equation.
  • Taking both coefficients from the wrong row. The intercept comes from the “(Intercept)” row. The slope comes from the row for the predictor. Do not use the predictor’s estimate as the intercept.
  • Reversing the predictor and response. A regression line predicts the response from the predictor used to fit it. Check the context and model description rather than assuming the row label tells you which variable is predicted.
  • Dropping a negative sign. Copy the sign of the predictor’s estimate exactly. A changed sign changes the direction and numerical value of a prediction.
  • Reporting the input as the prediction. The supplied value is \(x\), while the result of evaluating the line is \(\hat{y}\). State the predicted response, not the predictor value again.
  • Calling the prediction an actual result or a cause. As discussed in “Prediction Versus Observed Values” and “Regression Does Not Mean Causation in Slope Statements,” a line gives a prediction; individual observations can differ, and an association alone does not establish causation.

For a clear AP-style response, show the coefficient mapping, the substitution, and a sentence interpreting the calculated prediction. For example: “Using the intercept estimate 6.35 and distance coefficient 2.18, the predicted commute time for a distance of 8.5 kilometers is \(6.35+2.18(8.5)=24.88\) minutes.” This makes the source of each number and the meaning of the result clear.

Key takeaway: In simple regression output, use the “Estimate” in the “(Intercept)” row as \(a\), and the “Estimate” in the predictor’s row as \(b\). Write \(\hat{y}=a+bx\), substitute the requested predictor value, and describe the result as a predicted response in context.

Check Your Understanding

For each question, distinguish the coefficient estimates from other output values and state what the calculated number predicts.

  1. A model predicts delivery time in minutes from route distance in kilometers. The “Estimate” values are 8.2 for “(Intercept)” and 1.7 for “distance_km.” Write the equation and predict the delivery time for a 6-kilometer route.
  2. A coefficient table shows 15.0 in the intercept estimate column, 2.4 in the predictor estimate column, and 0.6 in the predictor standard-error column. Which two values belong in the prediction equation, and why?
  3. A model predicts plant height in centimeters from days since planting. The estimates are 4.8 for the intercept and 1.25 for “days.” Find the predicted height at 12 days.
  4. A regression of test time in seconds on practice minutes has an intercept estimate of 72 and a practice coefficient of \(-0.20\). Find the predicted test time for 30 practice minutes.
  5. Explain why a regression prediction should not automatically be described as the actual response for a particular individual.