Tutorials › AP Statistics › Least-Squares Mixed Practice Set

Least-squares regression · Tutorial 880 of 1000

Least-Squares Mixed Practice Set

Combine summary-statistic calculations, regression predictions, and quick checks that help catch errors before you interpret a fitted line.

Intermediate 9 min read

What You'll Learn

  • Calculate a regression slope from the correlation and sample standard deviations.
  • Find an intercept using the sample means and write the fitted equation.
  • Check a slope and intercept against the supplied summary statistics.
  • Make and verify a prediction using both the equation and the mean point.
  • Explain whether a prediction is interpolation or extrapolation.

One Regression Problem, Several Checks

In “Free-Response Practice Making a Regression Prediction,” you brought an equation, a prediction, and a justification together. This mixed practice set adds another task: when a problem supplies summary statistics, calculate the fitted line and check that your calculations agree with those statistics before predicting.

For a least-squares line \(\hat{y}=a+bx\), the earlier tutorials “Computing Slope From r and Summary Statistics” and “Finding the Line Through the Means” established how to calculate \(b\) and \(a\). Here, use those results as parts of a single workflow. A useful audit is to check that the slope agrees with \(r\) and the standard deviations, and that the line predicts \(\bar{y}\) when \(x=\bar{x}\). After that, verify a requested prediction with a second calculation.

Key idea: A summary-based regression calculation should fit together: \(b=r(s_y/s_x)\), \(a=\bar{y}-b\bar{x}\), and the fitted line predicts \(\bar{y}\) at \(x=\bar{x}\). Use these relationships to check the equation before interpreting or using it.

These checks can reveal arithmetic errors, a misplaced minus sign, or reversed predictor and response roles. They do not replace the data pattern when judging a prediction: as in “Predicting Within the Data Range,” compare the requested \(x\)-value with the observed range and consider whether the relationship is reasonably linear.

A Repeatable Calculation and Audit

Keep the roles of the variables fixed. The predictor \(x\) is used as the input, and the response \(y\) is what the line predicts. The slope uses the response standard deviation divided by the predictor standard deviation, so its units are response units per predictor unit.

1
Identify the summaries.
Write down \(\bar{x}\), \(s_x\), \(\bar{y}\), \(s_y\), and \(r\), along with what the variables measure.
2
Calculate and audit the slope.
Use \(b=r(s_y/s_x)\). Check that its sign matches \(r\) and that its units are response units per predictor unit.
3
Calculate and audit the intercept.
Use \(a=\bar{y}-b\bar{x}\), then check that \(a+b\bar{x}=\bar{y}\).
4
Predict and verify.
Substitute the requested predictor value into the equation. Check the result with \(\hat{y}=\bar{y}+b(x-\bar{x})\), an equivalent form based on the line passing through the means.
5
Interpret the prediction carefully.
Report the response in context and with units. Use the observed \(x\)-range and the described pattern to discuss how well supported the prediction is.

The mean-centered calculation in Step 4 is not a different regression line. Since \(a=\bar{y}-b\bar{x}\), substituting this into \(a+bx\) gives \(\bar{y}+b(x-\bar{x})\). It is a convenient independent check: instead of starting from the intercept, begin at the mean response and adjust by the slope times the distance from the mean predictor.

Worked Example: Fit and Check a Line From Summaries

Worked Example: Fit and Check a Line From Summaries

A fictional community energy project records daily sunlight duration and daily electricity output for a set of small solar installations. Let \(x\) be sunlight duration in hours and \(y\) be electricity output in kilowatt-hours. The summaries are \(\bar{x}=6\) hours, \(s_x=1.5\) hours, \(\bar{y}=24\) kilowatt-hours, \(s_y=5\) kilowatt-hours, and \(r=0.75\). The observed sunlight durations range from 3 to 9 hours, and the scatterplot is described as reasonably linear with no conspicuous outlier. Find the regression line and predict output for 7.2 hours of sunlight.

State. Sunlight duration is the predictor, measured in hours. Electricity output is the response, measured in kilowatt-hours.

Plan. Calculate the slope from \(r\), \(s_y\), and \(s_x\), then calculate the intercept from the means. Check that the line predicts the mean output at the mean sunlight duration. Predict at \(x=7.2\), verify the prediction in mean-centered form, and check the requested value against the observed range.

Do. Calculate the slope:

$$ b=r\left(\frac{s_y}{s_x}\right) =(0.75)\left(\frac{5}{1.5}\right) =(0.75)(3.3333) =2.5 $$

The slope is positive, as \(r\) is positive, and its units are kilowatt-hours per hour. As an arithmetic check, \(5/1.5=10/3\), and \(0.75(10/3)=2.5\). Now calculate the intercept:

$$ a=\bar{y}-b\bar{x} =24-(2.5)(6) =24-15 =9 $$

The equation is \(\hat{y}=9+2.5x\). Check the mean point: at \(x=6\), the line gives \(9+2.5(6)=9+15=24\) kilowatt-hours, which equals \(\bar{y}\). This agrees with the supplied means.

For 7.2 hours of sunlight:

$$ \hat{y}=9+2.5(7.2) =9+18 =27\text{ kilowatt-hours} $$

Verify using the mean-centered form:

$$ \hat{y}=24+2.5(7.2-6) =24+2.5(1.2) =24+3 =27\text{ kilowatt-hours} $$

Conclude. For an installation with 7.2 hours of sunlight, the fitted line predicts daily electricity output of 27 kilowatt-hours. This is interpolation because 7.2 hours is within the observed range of 3 to 9 hours. The described roughly linear pattern supports using the line for this prediction.

Worked Example: Audit a Reported Equation

Worked Example: Audit a Reported Equation

A fictional transportation class studies commute distance and commute time. Let \(x\) be distance in kilometers and \(y\) be time in minutes. The summaries are \(\bar{x}=12\) kilometers, \(s_x=4\) kilometers, \(\bar{y}=30\) minutes, \(s_y=8\) minutes, and \(r=0.50\). A student reports the line \(\hat{y}=24+0.5x\). Check the reported equation, find the correct line if needed, and predict commute time for a distance of 20 kilometers. The observed distances range from 4 to 22 kilometers.

State. Commute distance is the predictor in kilometers, and commute time is the response in minutes. The student’s equation has the correct variable order, but its coefficients still need to be checked.

Plan. Use the summary statistics to calculate the required slope and intercept. Then check whether the student’s line agrees with each result. Finally, predict at 20 kilometers using the corrected line and verify the calculation another way.

Do. The slope required by the summaries is:

$$ b=r\left(\frac{s_y}{s_x}\right) =(0.50)\left(\frac{8}{4}\right) =(0.50)(2) =1\text{ minute per kilometer} $$

The reported slope, 0.5 minute per kilometer, does not match. Calculate the intercept using the required slope:

$$ a=\bar{y}-b\bar{x} =30-(1)(12) =18\text{ minutes} $$

So the line consistent with the summaries is \(\hat{y}=18+x\). Check its mean-point prediction: \(18+1(12)=30\) minutes, matching \(\bar{y}\). The student’s line happens to pass through the mean point too, since \(24+0.5(12)=30\), but that check alone is not enough. Its slope fails the relationship \(b=r(s_y/s_x)\).

At \(x=20\) kilometers, the corrected line predicts:

$$ \hat{y}=18+1(20)=38\text{ minutes} $$

The mean-centered check gives the same result:

$$ \hat{y}=30+1(20-12) =30+8 =38\text{ minutes} $$

Conclude. The equation consistent with the given summaries is \(\hat{y}=18+x\), and it predicts a commute time of 38 minutes for a distance of 20 kilometers. Since 20 kilometers is within the observed range of 4 to 22 kilometers, this is interpolation. The range supports using the line for a prediction there, though the prompt would also need to provide information about the data pattern to assess linearity.

This example illustrates why a single check can be insufficient. The mean-point check tests whether the equation goes through \((\bar{x},\bar{y})\), but different slopes can pass through that same point. When the summaries are given, check both the slope formula and the mean point.

Worked Example: Correct a Sign and Check an Extrapolation

Worked Example: Correct a Sign and Check an Extrapolation

A fictional environmental club records the number of weekly litter-collection events and the mass of litter found along a trail. Let \(x\) be the number of events per week and \(y\) be litter mass in kilograms. The summaries are \(\bar{x}=10\) events, \(s_x=2\) events, \(\bar{y}=42\) kilograms, \(s_y=6\) kilograms, and \(r=-0.50\). A draft calculation gives the line \(\hat{y}=54+1.5x\). The observed event counts range from 6 to 14. Find the correct line and predict litter mass for 16 events per week.

State. Collection events per week are the predictor, and litter mass in kilograms is the response. Because \(r\) is negative, the slope should also be negative.

Plan. Calculate the slope and intercept directly from the summaries, then use the sign and mean-point checks to diagnose the draft equation. Calculate the requested prediction and verify it with the mean-centered form. Compare 16 with the observed range.

Do. The slope is:

$$ b=r\left(\frac{s_y}{s_x}\right) =(-0.50)\left(\frac{6}{2}\right) =(-0.50)(3) =-1.5\text{ kilograms per event} $$

The draft slope has the wrong sign. Calculate the intercept:

$$ a=\bar{y}-b\bar{x} =42-(-1.5)(10) =42+15 =57\text{ kilograms} $$

The corrected line is \(\hat{y}=57-1.5x\). Its mean-point check is \(57-1.5(10)=57-15=42\) kilograms, matching \(\bar{y}\). The draft line also fails that check: \(54+1.5(10)=69\), not 42.

At \(x=16\) events per week, the line gives:

$$ \hat{y}=57-1.5(16) =57-24 =33\text{ kilograms} $$

Using the mean-centered form verifies the result:

$$ \hat{y}=42-1.5(16-10) =42-1.5(6) =42-9 =33\text{ kilograms} $$

Conclude. The line consistent with the summaries is \(\hat{y}=57-1.5x\), and it predicts 33 kilograms of litter for 16 collection events per week. This is extrapolation because 16 is above the observed maximum of 14 events per week. The calculation can be made, but its prediction is less well supported because the data do not show whether the pattern continues beyond the observed range.

Common Mistakes and AP Exam Tips

Mixed questions can make it tempting to focus on the final number and skip the checks. A short audit often catches errors before they spread through the rest of the work.

  • Reversing the standard-deviation ratio. The slope is \(r(s_y/s_x)\), not \(r(s_x/s_y)\). Keep the response standard deviation on top so the result has response units per predictor unit.
  • Ignoring the sign. Since standard deviations are positive, the slope and \(r\) have the same sign when both variables vary. A reported slope with the opposite sign is a warning to recheck the calculation.
  • Checking only the mean point. A line can pass through \((\bar{x},\bar{y})\) and still have the wrong slope. Check the slope against \(r(s_y/s_x)\) as well.
  • Using the wrong mean in the intercept formula. The intercept is \(a=\bar{y}-b\bar{x}\). Subtract the slope times the predictor mean from the response mean.
  • Treating a prediction as an observed value. Write that the line predicts the response, and include the response units. A regression prediction is not a guarantee of an individual outcome.
  • Skipping the range check. Compare the requested predictor value with the observed \(x\)-range and identify interpolation or extrapolation. A value within the range does not by itself prove that the relationship is linear.
  • Rounding too early. Keep the supplied summaries and calculated coefficients at full available precision while working. Round the reported prediction sensibly, with units.

A complete response makes the reasoning visible: it identifies the predictor and response, shows the summary-based calculation, checks the fitted line, and gives the prediction in context. If a question provides no scatterplot or description of the pattern, do not claim that linearity has been established. Say what the available information does support, such as whether the requested value falls within the observed range.

Key takeaway: For summary-based practice, calculate \(b\), calculate \(a\), and check the line at the mean point. Then predict with the fitted equation, verify with the mean-centered form, and interpret the result using its response units and the observed predictor range.

Check Your Understanding

Use the supplied summaries and show enough work to make your calculations and checks clear.

  1. A study has \(\bar{x}=5\), \(s_x=2\), \(\bar{y}=18\), \(s_y=3\), and \(r=0.80\). Find \(b\) and \(a\), write the regression line, and check that it predicts \(\bar{y}\) at \(x=\bar{x}\).
  2. For the summaries in Question 1, predict \(y\) when \(x=7\). Verify the prediction using \(\bar{y}+b(x-\bar{x})\).
  3. A reported line has slope \(-2\), but the supplied correlation is \(0.40\) and both standard deviations are positive. What can you conclude before checking the intercept?
  4. A fitted line predicts a response at \(x=31\), while the observed predictor values range from 8 to 30. Identify the type of prediction and state one limitation to mention.
  5. Why does checking that a proposed line passes through \((\bar{x},\bar{y})\) not prove that its slope is correct?