Two Different Ideas: A Population Relationship and a Sample Line
In “Correlation Mixed Practice Set,” you practiced describing the linear association in observed data. A regression line takes that description a step further: it uses an explanatory variable \(x\) to predict a response variable \(y\). But a line fitted to sample data is not automatically the true relationship for an entire population. Keeping those ideas separate helps you describe predictions without claiming more than the data support.
The sample regression line is written \(\hat{y}=a+bx\). The values \(a\) and \(b\) are calculated from the sample. The symbol \(\hat{y}\), read “y-hat,” means the response value predicted by this fitted line for a specified \(x\). It is not the observed response \(y\) for a particular individual.
To describe a possible linear population relationship, we can use \(\mu_y=\alpha+\beta x\). Here, \(\mu_y\) is the population mean response for cases with explanatory-variable value \(x\), while \(\alpha\) and \(\beta\) describe the population line. Those population quantities are generally unknown. By contrast, \(a\) and \(b\) are calculated from the sample and give a fitted estimate of a linear pattern.
The population equation is a way to express a linear population relationship when such a relationship is reasonable. It does not mean that every individual with the same \(x\)-value has exactly the response \(\mu_y\). Individuals can differ, and their observed responses may fall above or below the population mean. Likewise, a sample point may fall above or below its predicted value from the fitted line.
The earlier tutorial “What the Correlation Coefficient Measures” introduced \(r\) as a summary of the direction and strength of the linear association in observed data. A regression line is also based on the sample, but its role is prediction: it gives a predicted response for each chosen \(x\). A strong correlation does not make the sample line identical to a population relationship, nor does it guarantee an accurate prediction for every individual.
What Does \(\hat{y}\) Represent?
For a particular \(x\), \(\hat{y}\) is the value on the fitted sample line. It is the line’s predicted response, expressed in the response variable’s units. It is not a promise about what will happen to every individual at that \(x\), and it is not necessarily the actual response of any one individual.
The distinction is familiar from residuals, introduced in “Strong Correlation Does Not Mean a Good Model.” For an observed case, the residual is \(y-\hat{y}\): the observed response minus the response predicted by the sample line. A positive residual means the observation is above the line; a negative residual means it is below. The difference between \(y\) and \(\hat{y}\) is one way to see that prediction and observation are not interchangeable.
A sample regression line is useful even though it is not the population relationship itself. It summarizes the linear pattern in the data at hand and can provide predictions. How well those predictions apply beyond the sample depends on the data, the form of the relationship, and the population to which the prediction is being applied. In particular, a straight line fitted to sample data does not prove that the population pattern is straight.
Worked Examples: Reading Predictions and Limits
Worked Example: A Line Fitted to Seedling Data
In a fictional sample, a gardener records daily light exposure, \(x\), in hours and seedling height, \(y\), in centimeters, for five seedlings. The paired data are \((1,10)\), \((2,12)\), \((3,13)\), \((4,16)\), and \((5,19)\). Find the sample regression line, then explain what it predicts when daily light exposure is 4 hours.
The means are \(\bar{x}=3\) hours and \(\bar{y}=14\) centimeters. For the least-squares line, the slope is the sum of the products of the centered values divided by the sum of squared centered \(x\)-values. The centered \(x\)-values are \(-2,-1,0,1,2\), and the centered \(y\)-values are \(-4,-2,-1,2,5\). Thus:
The intercept is \(a=\bar{y}-b\bar{x}\), so:
At \(x=4\) hours, the line predicts:
The predicted height is 16.2 centimeters. The observed height of the seedling that received 4 hours of light was 16 centimeters, so its residual is \(y-\hat{y}=16-16.2=-0.2\) centimeters. The fitted line’s prediction is close to this observed value, but the two values are not the same. Also, this five-seedling sample does not establish that the population mean height at 4 hours is exactly 16.2 centimeters.
Worked Example: A Predicted Response Is Not an Individual Result
Suppose a fictional sample of plants produces the fitted line \(\hat{y}=4.2+1.8x\), where \(x\) is daily light exposure in hours and \(y\) is plant height in centimeters. What response does the line predict for a plant receiving 6 hours of light? Does it say that every such plant will be that height?
Substitute \(x=6\) into the sample line:
The line predicts a height of 15.0 centimeters at 6 hours of light. This is the value on the fitted line, not a guarantee that a particular plant will measure 15.0 centimeters. If a plant in the sample receiving 6 hours measured 16 centimeters, its residual would be \(16-15=1\) centimeter; if it measured 13 centimeters, its residual would be \(13-15=-2\) centimeters. The line gives the same prediction at that \(x\), even when individual outcomes differ.
Nor does this calculation directly report the population mean height at 6 hours. The fitted line was calculated from a sample. It can be used as an estimate of a population pattern, but the predicted value remains a sample-based estimate.
Worked Example: Different Samples Can Give Different Lines
Imagine that two fictional samples are drawn from the same broad population of community gardens. For each garden, \(x\) is the number of weekly volunteer hours and \(y\) is the number of harvested produce boxes. Sample A has the points \((1,5)\), \((2,7)\), and \((3,9)\). Sample B has the points \((1,4)\), \((2,8)\), and \((3,10)\). Find each fitted line and explain why the lines need not match.
For Sample A, the means are \(\bar{x}=2\) and \(\bar{y}=7\). The centered \(x\)-values are \(-1,0,1\); the centered \(y\)-values are \(-2,0,2\). Therefore:
The intercept for Sample A is \(a_A=7-(2)(2)=3\), so its fitted line is \(\hat{y}=3+2x\).
For Sample B, \(\bar{x}=2\) and \(\bar{y}=22/3\). The centered \(x\)-values are again \(-1,0,1\), and the centered \(y\)-values are \(-10/3,2/3,8/3\). Thus:
The intercept is \(a_B=22/3-(3)(2)=4/3\), giving the fitted line \(\hat{y}=4/3+3x\). The lines differ: Sample A has slope 2, while Sample B has slope 3. Both lines summarize their own sample data. If the samples come from the same population, the difference between fitted lines illustrates that a sample line can vary from one sample to another. It does not, by itself, show that the population relationship has changed.
The population relationship is not directly observed just by fitting either of these sample lines. The samples provide information about the population, but their fitted lines remain sample results.
Worked Example: A Numerical Prediction Beyond the Sample Range
Return to the seedling line \(\hat{y}=7.4+2.2x\), which was fitted using light exposures from 1 to 5 hours. What does the line predict for a seedling receiving 8 hours? What limitation should accompany that calculation?
Substitution gives:
The line gives a predicted height of 25.0 centimeters at 8 hours. However, 8 hours is outside the range of light exposures in the sample, which ran from 1 to 5 hours. This is extrapolation: using a fitted line to predict at an \(x\)-value beyond the observed data. The arithmetic is straightforward, but the sample does not show whether the linear pattern continues to 8 hours. The population relationship could bend, level off, or behave differently there. Therefore, the prediction should not be presented as an established population outcome.
Common Mistakes and AP Exam Tips
- Calling \(\hat{y}\) the observed response. The predicted response comes from the line; \(y\) is the actual recorded response. For an observed case, use the residual \(y-\hat{y}\) to describe the difference.
- Calling a sample line the true population line. The coefficients \(a\) and \(b\) are calculated from sample data. Say “the fitted sample line predicts” rather than claiming that it exactly gives the population relationship.
- Claiming the prediction applies to every individual. A line gives one predicted response at a chosen \(x\). Individuals with that same \(x\) can have different observed responses.
- Assuming the population relationship must be linear. A fitted straight line summarizes a linear pattern in the sample. The line alone does not prove that the population pattern is straight; inspect the scatterplot for curvature or other structure.
- Ignoring the range of the observed \(x\)-values. A prediction outside that range is extrapolation. Identify it and explain that the data do not show whether the pattern continues there.
For a full-credit response, state that \(\hat{y}\) is the predicted response from the sample regression line, give its value in context and units, and avoid treating it as a guaranteed individual outcome or a known population truth. If the question asks about a specific observed case, distinguish its actual \(y\) from its predicted \(\hat{y}\).
Check Your Understanding
Use the distinction between a population relationship, a sample line, and an observed response to answer these questions.
- A fitted sample line is \(\hat{y}=6+1.5x\). Find the predicted response at \(x=8\), and explain what the prediction represents.
- At \(x=8\), an individual’s observed response is 20. Using the line in Question 1, calculate the residual and explain its sign.
- Explain why two different samples from the same population could produce different fitted regression lines.
- A student says, “The fitted line predicts 12, so every individual with this \(x\)-value will have a response of 12.” Identify the error and rewrite the statement accurately.
- A line fitted using \(x\)-values from 2 to 10 is used to predict at \(x=15\). Name the issue and explain one reason the prediction may be unreliable.