Tutorials › AP Statistics › Planning a Complete Regression Analysis

Comparing and communicating regression models · Tutorial 981 of 1000

Planning a Complete Regression Analysis

Learn a repeatable way to organize a regression analysis, check whether a linear model is useful, and communicate what the results do—and do not—show.

Intermediate 11 min read

What You'll Learn

  • Turn a real-world question into clearly identified explanatory and response variables with units.
  • Use a scatterplot to assess direction, form, strength, and unusual observations before fitting a line.
  • Organize the regression equation, \(r\), \(r^2\), and residual standard deviation as complementary summaries.
  • Use residuals and a residual plot to judge whether a straight-line model is a reasonable description.
  • Write a conclusion that stays within the data’s range and avoids unsupported causal claims.

From a Question to a Defensible Conclusion

A regression analysis is more than calculating a line. It begins with a question about two quantitative variables and ends with a conclusion that explains what the data show, how well a linear model describes them, and what the model cannot establish. A clear order helps you avoid fitting a line before you have checked whether a line makes sense.

In “Model Fit Diagnosis Mixed Practice Set,” you practiced separating evidence from scatterplots, residuals, and regression summaries. Here, we organize those tools into a complete analysis. The new technique is to plan the analysis as a chain: each stage answers a different question, and the conclusion draws only on evidence that the earlier stages support.

Analysis map: State the question and identify the cases and variables; inspect a scatterplot; fit and describe a linear model if appropriate; check residuals and unusual observations; then answer the original question in context, with any limitations.

This is a descriptive analysis of a relationship in observed data, not an inference procedure with a single set of formal conditions. Still, choices about the cases, the graph, and the fit matter. A good analysis does not treat a high \(r\) or \(r^2\) as proof that the line is appropriate, and it does not treat a fitted line as proof that changes in \(x\) cause changes in \(y\).

Plan Before You Calculate

Start by translating the question into statistical roles. The explanatory variable, \(x\), is used to explain or predict the response variable, \(y\). State what one case represents, name each variable, and include its units. The way the question is phrased may suggest these roles, but the roles should also make sense in context.

Next, make a scatterplot with \(x\) on the horizontal axis and \(y\) on the vertical axis. Describe the overall direction, form, and strength of the association, and note any unusual observations. As discussed in “Residual Plot Versus Scatterplot Diagnosis,” this graph shows the original data; it is not the same as a residual plot. If the scatterplot has a clear bend or another serious departure from a roughly straight pattern, pause before using a linear model.

If a line is reasonable to consider, describe the fitted equation \(\hat{y}=a+bx\). Interpret the slope \(b\) as the predicted change in \(y\) for an increase of one unit in \(x\), in the context and with units. The intercept \(a\) is the predicted response when \(x=0\); interpret it only if zero is meaningful and relevant to the data.

Use the regression summaries for different purposes. The correlation \(r\) describes the direction and strength of the linear association. The coefficient of determination, \(r^2\), is the proportion of variation in the observed response values accounted for by the linear model. The residual standard deviation, \(s\), describes the typical size of residuals in response-variable units. None of these replaces looking at the scatterplot and residual plot.

1
Frame the question.
Identify the cases, the explanatory variable, the response variable, their units, and the purpose of the analysis.
2
Inspect the scatterplot.
Describe direction, form, strength, and unusual points. Decide whether a straight-line model is sensible to examine.
3
Fit and summarize the line.
Report the equation and relevant regression summaries. Interpret them in context rather than listing numbers alone.
4
Check errors and conclude.
Use residuals and a residual plot to assess the fit. Answer the original question, while respecting the data’s scope and limits.

Worked Example: A Complete Analysis of Light and Seedling Height

Worked Example: Do More Hours of Light Go with Taller Seedlings?

Original AP-style question. In an invented classroom experiment, each case is a seedling of the same variety measured at the same age. The explanatory variable \(x\) is the number of hours of light per day, and the response variable \(y\) is height in centimeters. The observed pairs are \((1,13),(3,15),(5,20),(7,24),(9,27),(11,33)\). Describe the relationship and decide whether a linear model is useful for these observations.

State. We want to describe how daily hours of light are associated with seedling height in this set of six observations. The explanatory variable is hours of light per day, and the response variable is height in centimeters. The data can describe an association, but this invented example does not by itself establish a causal effect.

Plan. First inspect the scatterplot for direction, form, strength, and unusual points. If the pattern is roughly linear, calculate the least-squares regression line and summarize its fit. Then calculate residuals and examine their pattern before deciding what conclusion is justified.

Do. A scatterplot of the six pairs would show a strong positive, roughly linear association: seedlings receiving more hours of light tend to be taller. There is no obvious isolated point or pronounced bend in this small set. This is a reason to examine a linear model, not a guarantee that the model fits perfectly.

The means are \(\bar{x}=6\) hours per day and \(\bar{y}=22\) centimeters. The centered sums are \(S_{xx}=70\), \(S_{xy}=140\), and \(S_{yy}=284\), where \(S_{xx}\) summarizes variation in \(x\), \(S_{xy}\) summarizes the joint variation, and \(S_{yy}\) summarizes variation in \(y\). The least-squares slope is \(b=S_{xy}/S_{xx}\), and the intercept is \(a=\bar{y}-b\bar{x}\):

$$ b=\frac{140}{70}=2 \quad\text{centimeters per additional hour of light per day}, \qquad a=22-(2)(6)=10\text{ centimeters} $$

The fitted line is \(\hat{y}=10+2x\). Within this data set, the line predicts that height increases by 2 centimeters for each additional hour of light per day. The intercept says the fitted line predicts a height of 10 centimeters at zero hours of light per day; that value is outside the observed \(x\)-range of 1 to 11 hours and may not be meaningful for this setting, so it should not be given a practical interpretation.

The data’s \(r^2\) and \(r\) are

$$ r^2=\frac{S_{xy}^2}{S_{xx}S_{yy}} =\frac{140^2}{(70)(284)} =\frac{490}{497} \approx0.9859, \qquad r=\sqrt{r^2}\approx0.9929 $$

The correlation is positive because the slope is positive. In these observations, about 98.59% of the variation in seedling heights is accounted for by the linear model relating height to hours of light per day. This is a description of how closely the model accounts for variation in these data, not a claim that light explains that percentage of height differences in every setting.

To check the fit, use \(\hat{y}=10+2x\) to calculate each predicted height and subtract it from the observed height. The residuals, in centimeters, are \(1,-1,0,0,-1,1\), respectively. Their squared values sum to \(SSE=4\). This agrees with the calculation \(SSE=S_{yy}-S_{xy}^2/S_{xx}=284-140^2/70=4\). Thus, the residual standard deviation is

$$ s=\sqrt{\frac{SSE}{n-2}} =\sqrt{\frac{4}{6-2}} =1\text{ centimeter} $$

The residual plot would show residuals close to zero, with no clear curve or fan shape. The residuals are not all zero, but their pattern does not show a strong systematic failure of the line. With only six observations, this check is limited; it should not be described as proof that the model will work in other settings.

Conclude. The scatterplot and residuals support using a straight line to describe these six invented observations. The fitted line predicts a 2-centimeter increase in height for each additional hour of light per day, and the residuals are typically about 1 centimeter from the line. These results describe an association among the observed seedlings. They do not establish that extra light caused the height differences, and the model should not be extended beyond the observed hours without considering the risks of extrapolation discussed in “What Extrapolation Means.”

When the First Graph Warns You About a Line

A planned analysis also needs a decision point: sometimes the scatterplot suggests that a straight line does not describe the relationship well. The next example shows why a large \(r\) is not enough to skip the residual check.

Worked Example: A Strong Correlation with Curved Residuals

Original AP-style question. In a separate invented exercise, \(x\) is the number of practice sessions and \(y\) is a skill score. The observations are \((1,4),(2,7),(3,12),(4,19),(5,28)\). Assess whether a linear model is an adequate description.

Solution. The scatterplot rises, but the increases in score become larger as sessions increase, giving the pattern a noticeable upward bend. To see what a straight-line fit would miss, calculate its equation. Here, \(\bar{x}=3\), \(\bar{y}=14\), \(S_{xx}=10\), \(S_{xy}=60\), and \(S_{yy}=374\). Therefore

$$ b=\frac{60}{10}=6, \qquad a=14-(6)(3)=-4, \qquad \hat{y}=-4+6x $$

The correlation is positive and high:

$$ r^2=\frac{60^2}{(10)(374)} =\frac{360}{374} \approx0.9626, \qquad r\approx0.9811 $$

But the fitted values for \(x=1,2,3,4,5\) are \(2,8,14,20,26\). Subtracting these from the observed scores gives residuals \(2,-1,-2,-1,2\) score points. They are positive at both ends and negative in the middle, a curved pattern. Also, \(SSE=374-60^2/10=14\), so \(s=\sqrt{14/(5-2)}\approx2.1602\) score points.

Conclusion. Although \(r\approx0.9811\) describes a strong positive linear association, the residuals show a systematic curve. The high correlation does not make the straight-line model adequate; it misses the changing pattern in scores across practice sessions. The residual evidence should be part of the conclusion, not omitted because \(r\) is close to 1.

Use the Model to Answer a Question Carefully

When the scatterplot and residual plot support a line, the fitted equation can help answer a prediction question. A complete answer still identifies the requested \(x\)-value, calculates the prediction, checks whether that value is within the observed range, and explains what the prediction represents.

Worked Example: Predicting a Score Within the Observed Range

Original AP-style question. In an invented study group, \(x\) is hours spent reviewing and \(y\) is a practice-test score in points. The five observations are \((1,62),(2,66),(3,71),(4,73),(5,78)\). Use a linear model to predict the score for 4.5 hours of review, and describe the strength and limitations of that prediction.

Solution. The values increase in a roughly straight pattern. The means are \(\bar{x}=3\) hours and \(\bar{y}=70\) points. The centered sums are \(S_{xx}=10\), \(S_{xy}=39\), and \(S_{yy}=154\). The least-squares line is

$$ b=\frac{39}{10}=3.9\text{ points per hour}, \qquad a=70-(3.9)(3)=58.3\text{ points}, \qquad \hat{y}=58.3+3.9x $$

The model predicts a score of

$$ \hat{y}=58.3+(3.9)(4.5)=75.85\text{ points} $$

The observed review-hour range is 1 to 5, so 4.5 hours is within the range. As explained in “Interpolation Versus Extrapolation,” this is interpolation. The residual standard deviation is based on \(SSE=154-39^2/10=1.9\):

$$ s=\sqrt{\frac{1.9}{5-2}} \approx0.796\text{ points} $$

The correlation is positive, with \(r^2=39^2/((10)(154))=1521/1540\approx0.9877\), so \(r\approx0.9938\). The fitted line predicts about 75.85 points at 4.5 hours. This is a model prediction for the response at that review time, not a guarantee of the score for any particular student. The invented observations also do not establish that more review time causes higher scores.

Common Mistakes and What a Strong Conclusion Says

  • Starting with the equation instead of the question. Without identifying the cases, variables, and units, it is easy to reverse \(x\) and \(y\) or give an uninterpretable slope. State what one case represents and name both variables first.
  • Calling a relationship linear from \(r\) alone. A strong correlation can occur when the scatterplot bends. Describe the scatterplot and inspect the residual plot before concluding that a line is suitable.
  • Reporting summaries without interpretation. A response that lists \(r^2\) or \(s\) without explaining its meaning is incomplete. State the proportion of response variation accounted for by the linear model, or describe the typical residual size in response units.
  • Interpreting the intercept automatically. The intercept is the predicted response at \(x=0\), but zero may be outside the observed range or nonsensical in context. Explain that limitation rather than assigning the intercept an unsupported practical meaning.
  • Treating a prediction as certain or causal. A fitted value is a model-based prediction, not an exact outcome. Association alone does not show causation; consider how the data were collected before making causal claims.
  • Forgetting the range and scope. Check the observed \(x\)-range before making a prediction, as emphasized in “Reliability Within the Range of Data.” An in-range prediction is generally better supported than an extrapolation, but it is not automatically reliable in every population or setting.
  • Using “the model fits” as an absolute claim. A residual plot without an obvious pattern supports the line as a useful description of these data. It does not prove the relationship is exactly linear or guarantee performance beyond the observed cases.

For full-credit communication, connect evidence to the claim: describe the scatterplot, state and interpret the line, use residuals to assess fit, and answer the question in context. Include units and qualify the conclusion when the data or design limit what can be said.

Key takeaway: Plan the regression analysis from the question outward. Let the scatterplot guide whether to fit a line, let residuals check what the line misses, and let the conclusion reflect the data’s context, range, and limitations.

Check Your Understanding

Use the analysis sequence from this tutorial to answer each question.

  1. A study records each employee’s weekly training hours and a performance rating. Which variable would you place on the horizontal axis if the goal is to predict rating from training hours? Name the explanatory and response variables.
  2. A scatterplot has a strong positive association but bends upward. What should you check before deciding that a linear model is adequate?
  3. A fitted line has slope 4.2 points per hour. Write a sentence interpreting the slope in context, including the predicted nature of the change.
  4. A regression output reports \(r^2=0.81\). What does this say about the variation in the response for the data used to fit the model?
  5. The residual plot shows residuals that are positive at low and high \(x\)-values and negative in the middle. What does this pattern suggest about the linear fit?
  6. A requested prediction uses an \(x\)-value beyond the observed range. What check should you mention, and why should the conclusion be qualified?