Tutorials › AP Statistics › Population Model Versus Sample Regression Line

Linear regression models · Tutorial 841 of 1000

Population Model Versus Sample Regression Line

Distinguish a sample’s fitted regression line from a population relationship, and interpret what its predicted response says—and does not say—about an individual.

Intermediate 9 min read

What You'll Learn

  • Distinguish a sample regression line from a population relationship.
  • Explain what the population mean response at a given explanatory-variable value represents.
  • Interpret \(\hat{y}\) as a prediction from a fitted sample line.
  • Separate a predicted response from an individual’s observed response.
  • Explain why different samples can produce different fitted lines.
  • Recognize why a fitted line alone does not establish the true form of a population relationship.

Two Different Ideas: A Population Relationship and a Sample Line

In “Correlation Mixed Practice Set,” you practiced describing the linear association in observed data. A regression line takes that description a step further: it uses an explanatory variable \(x\) to predict a response variable \(y\). But a line fitted to sample data is not automatically the true relationship for an entire population. Keeping those ideas separate helps you describe predictions without claiming more than the data support.

The sample regression line is written \(\hat{y}=a+bx\). The values \(a\) and \(b\) are calculated from the sample. The symbol \(\hat{y}\), read “y-hat,” means the response value predicted by this fitted line for a specified \(x\). It is not the observed response \(y\) for a particular individual.

Definition: The sample regression line is a line fitted to the observed sample data. For an explanatory-variable value \(x\), its predicted response is \(\hat{y}=a+bx\). A population relationship describes how the response tends to vary across the entire population; if its mean response follows a straight-line pattern, that population line is distinct from the line estimated from any one sample.

To describe a possible linear population relationship, we can use \(\mu_y=\alpha+\beta x\). Here, \(\mu_y\) is the population mean response for cases with explanatory-variable value \(x\), while \(\alpha\) and \(\beta\) describe the population line. Those population quantities are generally unknown. By contrast, \(a\) and \(b\) are calculated from the sample and give a fitted estimate of a linear pattern.

$$ \text{Population mean relationship: }\mu_y=\alpha+\beta x \qquad \text{Sample regression line: }\hat{y}=a+bx $$

The population equation is a way to express a linear population relationship when such a relationship is reasonable. It does not mean that every individual with the same \(x\)-value has exactly the response \(\mu_y\). Individuals can differ, and their observed responses may fall above or below the population mean. Likewise, a sample point may fall above or below its predicted value from the fitted line.

The earlier tutorial “What the Correlation Coefficient Measures” introduced \(r\) as a summary of the direction and strength of the linear association in observed data. A regression line is also based on the sample, but its role is prediction: it gives a predicted response for each chosen \(x\). A strong correlation does not make the sample line identical to a population relationship, nor does it guarantee an accurate prediction for every individual.

What Does \(\hat{y}\) Represent?

For a particular \(x\), \(\hat{y}\) is the value on the fitted sample line. It is the line’s predicted response, expressed in the response variable’s units. It is not a promise about what will happen to every individual at that \(x\), and it is not necessarily the actual response of any one individual.

The distinction is familiar from residuals, introduced in “Strong Correlation Does Not Mean a Good Model.” For an observed case, the residual is \(y-\hat{y}\): the observed response minus the response predicted by the sample line. A positive residual means the observation is above the line; a negative residual means it is below. The difference between \(y\) and \(\hat{y}\) is one way to see that prediction and observation are not interchangeable.

Key distinction: \(y\) is an observed response for an individual. \(\hat{y}\) is the response predicted by the fitted sample line at that individual’s \(x\)-value. A prediction describes the line’s estimate, not a guaranteed individual outcome or a known population truth.

A sample regression line is useful even though it is not the population relationship itself. It summarizes the linear pattern in the data at hand and can provide predictions. How well those predictions apply beyond the sample depends on the data, the form of the relationship, and the population to which the prediction is being applied. In particular, a straight line fitted to sample data does not prove that the population pattern is straight.

Worked Examples: Reading Predictions and Limits

Worked Example: A Line Fitted to Seedling Data

In a fictional sample, a gardener records daily light exposure, \(x\), in hours and seedling height, \(y\), in centimeters, for five seedlings. The paired data are \((1,10)\), \((2,12)\), \((3,13)\), \((4,16)\), and \((5,19)\). Find the sample regression line, then explain what it predicts when daily light exposure is 4 hours.

The means are \(\bar{x}=3\) hours and \(\bar{y}=14\) centimeters. For the least-squares line, the slope is the sum of the products of the centered values divided by the sum of squared centered \(x\)-values. The centered \(x\)-values are \(-2,-1,0,1,2\), and the centered \(y\)-values are \(-4,-2,-1,2,5\). Thus:

$$ b=\frac{(-2)(-4)+(-1)(-2)+(0)(-1)+(1)(2)+(2)(5)} {(-2)^2+(-1)^2+0^2+1^2+2^2} =\frac{22}{10}=2.2 $$

The intercept is \(a=\bar{y}-b\bar{x}\), so:

$$ a=14-(2.2)(3)=7.4 \qquad \hat{y}=7.4+2.2x $$

At \(x=4\) hours, the line predicts:

$$ \hat{y}=7.4+2.2(4)=16.2\text{ centimeters} $$

The predicted height is 16.2 centimeters. The observed height of the seedling that received 4 hours of light was 16 centimeters, so its residual is \(y-\hat{y}=16-16.2=-0.2\) centimeters. The fitted line’s prediction is close to this observed value, but the two values are not the same. Also, this five-seedling sample does not establish that the population mean height at 4 hours is exactly 16.2 centimeters.

Worked Example: A Predicted Response Is Not an Individual Result

Suppose a fictional sample of plants produces the fitted line \(\hat{y}=4.2+1.8x\), where \(x\) is daily light exposure in hours and \(y\) is plant height in centimeters. What response does the line predict for a plant receiving 6 hours of light? Does it say that every such plant will be that height?

Substitute \(x=6\) into the sample line:

$$ \hat{y}=4.2+1.8(6)=4.2+10.8=15.0\text{ centimeters} $$

The line predicts a height of 15.0 centimeters at 6 hours of light. This is the value on the fitted line, not a guarantee that a particular plant will measure 15.0 centimeters. If a plant in the sample receiving 6 hours measured 16 centimeters, its residual would be \(16-15=1\) centimeter; if it measured 13 centimeters, its residual would be \(13-15=-2\) centimeters. The line gives the same prediction at that \(x\), even when individual outcomes differ.

Nor does this calculation directly report the population mean height at 6 hours. The fitted line was calculated from a sample. It can be used as an estimate of a population pattern, but the predicted value remains a sample-based estimate.

Worked Example: Different Samples Can Give Different Lines

Imagine that two fictional samples are drawn from the same broad population of community gardens. For each garden, \(x\) is the number of weekly volunteer hours and \(y\) is the number of harvested produce boxes. Sample A has the points \((1,5)\), \((2,7)\), and \((3,9)\). Sample B has the points \((1,4)\), \((2,8)\), and \((3,10)\). Find each fitted line and explain why the lines need not match.

For Sample A, the means are \(\bar{x}=2\) and \(\bar{y}=7\). The centered \(x\)-values are \(-1,0,1\); the centered \(y\)-values are \(-2,0,2\). Therefore:

$$ b_A=\frac{(-1)(-2)+(0)(0)+(1)(2)}{(-1)^2+0^2+1^2} =\frac{4}{2}=2 $$

The intercept for Sample A is \(a_A=7-(2)(2)=3\), so its fitted line is \(\hat{y}=3+2x\).

For Sample B, \(\bar{x}=2\) and \(\bar{y}=22/3\). The centered \(x\)-values are again \(-1,0,1\), and the centered \(y\)-values are \(-10/3,2/3,8/3\). Thus:

$$ b_B=\frac{(-1)(-10/3)+(0)(2/3)+(1)(8/3)}{2} =\frac{6}{2}=3 $$

The intercept is \(a_B=22/3-(3)(2)=4/3\), giving the fitted line \(\hat{y}=4/3+3x\). The lines differ: Sample A has slope 2, while Sample B has slope 3. Both lines summarize their own sample data. If the samples come from the same population, the difference between fitted lines illustrates that a sample line can vary from one sample to another. It does not, by itself, show that the population relationship has changed.

The population relationship is not directly observed just by fitting either of these sample lines. The samples provide information about the population, but their fitted lines remain sample results.

Worked Example: A Numerical Prediction Beyond the Sample Range

Return to the seedling line \(\hat{y}=7.4+2.2x\), which was fitted using light exposures from 1 to 5 hours. What does the line predict for a seedling receiving 8 hours? What limitation should accompany that calculation?

Substitution gives:

$$ \hat{y}=7.4+2.2(8)=7.4+17.6=25.0\text{ centimeters} $$

The line gives a predicted height of 25.0 centimeters at 8 hours. However, 8 hours is outside the range of light exposures in the sample, which ran from 1 to 5 hours. This is extrapolation: using a fitted line to predict at an \(x\)-value beyond the observed data. The arithmetic is straightforward, but the sample does not show whether the linear pattern continues to 8 hours. The population relationship could bend, level off, or behave differently there. Therefore, the prediction should not be presented as an established population outcome.

Common Mistakes and AP Exam Tips

  • Calling \(\hat{y}\) the observed response. The predicted response comes from the line; \(y\) is the actual recorded response. For an observed case, use the residual \(y-\hat{y}\) to describe the difference.
  • Calling a sample line the true population line. The coefficients \(a\) and \(b\) are calculated from sample data. Say “the fitted sample line predicts” rather than claiming that it exactly gives the population relationship.
  • Claiming the prediction applies to every individual. A line gives one predicted response at a chosen \(x\). Individuals with that same \(x\) can have different observed responses.
  • Assuming the population relationship must be linear. A fitted straight line summarizes a linear pattern in the sample. The line alone does not prove that the population pattern is straight; inspect the scatterplot for curvature or other structure.
  • Ignoring the range of the observed \(x\)-values. A prediction outside that range is extrapolation. Identify it and explain that the data do not show whether the pattern continues there.

For a full-credit response, state that \(\hat{y}\) is the predicted response from the sample regression line, give its value in context and units, and avoid treating it as a guaranteed individual outcome or a known population truth. If the question asks about a specific observed case, distinguish its actual \(y\) from its predicted \(\hat{y}\).

Key takeaway: The sample line \(\hat{y}=a+bx\) is fitted to observed data; the population relationship is the broader pattern the sample may help estimate. At a chosen \(x\), \(\hat{y}\) is a sample-line prediction, not a guaranteed individual response or proof of the population truth.

Check Your Understanding

Use the distinction between a population relationship, a sample line, and an observed response to answer these questions.

  1. A fitted sample line is \(\hat{y}=6+1.5x\). Find the predicted response at \(x=8\), and explain what the prediction represents.
  2. At \(x=8\), an individual’s observed response is 20. Using the line in Question 1, calculate the residual and explain its sign.
  3. Explain why two different samples from the same population could produce different fitted regression lines.
  4. A student says, “The fitted line predicts 12, so every individual with this \(x\)-value will have a response of 12.” Identify the error and rewrite the statement accurately.
  5. A line fitted using \(x\)-values from 2 to 10 is used to predict at \(x=15\). Name the issue and explain one reason the prediction may be unreliable.