Tutorials › AP Statistics › Free-Response Practice on Regression in Context

Regression and context · Tutorial 939 of 1000

Free-Response Practice on Regression in Context

Learn to write a complete regression response that connects the model’s slope to the study context and evaluates whether the design justifies a causal conclusion.

Intermediate 9 min read

What You'll Learn

  • Interpret a regression slope using the explanatory and response variables, their units, and the observed cases
  • Distinguish an observational study from a randomized experiment when evaluating causation
  • Explain how a specific possible confounder could contribute to an observed association
  • Organize a response so each part directly answers the question asked
  • Use scoring notes to identify what makes an explanation complete

Build a Complete Response from the Model and the Study

A regression free-response question may ask you to interpret a slope and then decide whether the relationship described by the model is causal. These are connected tasks, but they are not the same task. The slope describes a predicted change in the response for a change in the explanatory variable. The study design determines whether it is reasonable to say that changing the explanatory variable causes a change in the response.

In “Keeping Context in Every Regression Sentence,” we practiced naming the cases, variables, and units when interpreting a regression result. In “Association Versus Causation in Regression,” we distinguished a recorded association from a causal effect. Here, the new skill is to combine those ideas in a concise, organized free-response answer—and to explain the design reasoning rather than relying on the phrase “correlation does not imply causation.”

Key strategy: A complete answer connects the model to the cases and units, identifies how the explanatory variable was studied, and states what that design does or does not allow you to conclude about causation.

A useful response plan has four parts. You do not need to label these parts in every answer, but checking them before you finish can help prevent missing a scoring element.

1
Identify what the question asks.
Separate the requested model interpretation from the requested conclusion about causation.
2
Interpret the regression result in context.
Name the cases, explanatory variable, response variable, and units. State that the fitted line predicts a change in the response for a specified change in the explanatory variable.
3
Describe the study design.
Say whether researchers recorded the variables or deliberately assigned a treatment. Do not infer the design from the fitted line.
4
Explain the causal conclusion.
State whether the design supports a causal claim. For observational data, name a plausible third variable and explain how it could be related to both variables.

The explanation of a possible confounder should be specific. Naming “other factors” is usually not enough: identify one plausible variable and describe its connection to both the explanatory variable and the response. The variable should make sense in the setting, but you should present it as a possible explanation unless the study actually measured and established its role.

Worked Free-Response Questions

Worked Example: Sunlight and Tomato Harvest

Fictional AP-style question. A community garden coordinator records data from 48 tomato plots at several gardens. For each plot, the coordinator records the average number of hours of direct sunlight per day and the mass of tomatoes harvested during one week, in kilograms. The plots were not randomly selected, and no sunlight exposure was assigned. A least-squares regression line for predicting weekly harvest from daily sunlight is \(\hat{y}=1.6+0.42x\). A gardener claims that giving a plot one extra hour of sunlight will cause its weekly tomato harvest to increase by \(0.42\) kilograms.

(a) Interpret the slope of the regression line in context. (b) Does this study provide convincing evidence that an extra hour of sunlight causes a plot’s weekly harvest to increase? Explain.

State and plan. The cases are tomato plots. The explanatory variable \(x\) is average daily direct sunlight in hours, and the response variable \(y\) is weekly tomato harvest in kilograms. The prompt says sunlight was recorded rather than assigned, so this is an observational study. Interpret the slope as a model prediction, then use the design to evaluate the causal claim.

Do (part a). The slope is \(0.42\) kilograms per hour. A complete interpretation is: Among the 48 tomato plots studied, the fitted line predicts that plots receiving one additional hour of average direct sunlight per day will have a weekly tomato harvest that is 0.42 kilograms greater, on average. The units matter: the predictor changes by one hour per day, and the predicted response difference is in kilograms per week.

Do (part b). The study does not establish that an extra hour of sunlight causes a plot’s harvest to increase. Researchers observed existing sunlight and harvest; they did not randomly assign plots to different sunlight exposures. For example, gardeners may water plots with more sunlight differently, and watering could be related to both sunlight exposure and harvest. Soil quality is another plausible factor: it may vary among plots and relate to both where plots receive sunlight and how much they produce. These possibilities show why the observed relationship alone does not isolate a causal effect.

Conclude. The regression supports an association between daily sunlight and weekly harvest for the plots studied, and its slope gives a predicted difference in context. It does not show that increasing one plot’s sunlight would cause its harvest to rise by exactly \(0.42\) kilograms per week.

Illustrative scoring notes (4 points, not an official rubric): One point for identifying the response and explanatory variables with their units; one for correctly interpreting \(0.42\) as a predicted response change per additional hour per day; one for identifying the study as observational because sunlight was recorded, not assigned; and one for explaining that causation is not established, supported by a plausible third variable linked to both sunlight and harvest. “The relationship is not causal because correlation is not causation” would not earn the explanation point by itself.

Worked Example: Practice Time and a Skills Assessment

Fictional AP-style question. A coach records the number of minutes that 72 volunteer swimmers practiced a particular turn technique during the week before a skills assessment. The coach also records each swimmer’s assessment score, in points. Swimmers choose their own practice time; the coach does not assign it. The fitted line for predicting score from practice time is \(\hat{y}=61+0.08x\), where \(x\) is minutes of practice and \(y\) is the assessment score.

(a) Interpret the slope. (b) A swimmer says, “If I practice 30 extra minutes, my score will go up by 2.4 points.” Explain what the model does and does not support.

State and plan. The cases are the 72 volunteer swimmers. Practice time is the explanatory variable, measured in minutes, and assessment score is the response, measured in points. Since each swimmer chose how much to practice, the data are observational. Interpret the slope first, then distinguish a predicted comparison across swimmers from a guaranteed change for one swimmer.

Do (part a). The slope is \(0.08\) points per minute. Among these swimmers, the fitted line predicts an assessment score \(0.08\) points higher, on average, for each additional minute of technique practice during the week. For a 30-minute difference, the model’s predicted score difference is

$$ 0.08(30)=2.4\text{ points}. $$

So the model predicts a 2.4-point difference between swimmers whose practice times differ by 30 minutes, within the range and setting represented by the data. It does not say that every swimmer will gain 2.4 points by personally adding 30 minutes of practice.

Do and conclude (part b). The calculation matches the fitted line, but the swimmer’s statement turns an observed association into a guaranteed individual outcome. Because the coach recorded practice time rather than randomly assigning it, the study does not establish that extra practice causes a score increase. Prior skill is one plausible third variable: swimmers with stronger technique may choose to practice more and may also earn higher assessment scores. Motivation could also be related to both practice time and performance. These are possible explanations, not proven causes.

A careful conclusion is: Among the volunteer swimmers studied, more technique-practice time was associated with higher assessment scores. The fitted line predicts a 2.4-point score difference for swimmers whose practice times differ by 30 minutes, but this observational study does not show that an individual swimmer will gain 2.4 points by adding 30 minutes.

Illustrative scoring notes (4 points, not an official rubric): One point for the slope interpretation with points and minutes; one for calculating and contextualizing \(0.08(30)=2.4\); one for explaining that swimmers chose their own practice time, so this is observational; and one for rejecting the guaranteed causal claim with a relevant explanation, such as prior skill being related to both practice and score. Saying only “it depends on the swimmer” is too vague to explain the study-design limitation.

Worked Example: Randomly Assigned Reminder Messages

Fictional AP-style question. Researchers recruit 60 volunteer adult learners for a two-week study of an online course. They randomly assign 30 learners to receive a daily study reminder message and 30 to receive the course’s standard weekly message. At the end of the study, researchers record each learner’s number of completed lessons. Let \(x=0\) represent the standard weekly message and \(x=1\) represent the daily reminder. The group means are 7.1 completed lessons for \(x=0\) and 8.5 for \(x=1\). A fitted regression line is \(\hat{y}=7.1+1.4x\).

(a) Interpret the slope. (b) Does the study support a causal conclusion about the effect of assignment to daily reminder messages? Explain the role of random assignment and state an appropriate limit on the conclusion.

State and plan. The cases are the 60 volunteer adult learners. The explanatory variable is assigned message type, coded 0 or 1, and the response is the number of lessons completed in two weeks. Unlike the earlier observational examples, this is an experiment because researchers assigned the message condition at random. Explain the fitted group-mean difference and then assess what random assignment supports.

Do (part a). The slope is \(1.4\) lessons per one-unit increase in \(x\). Since \(x\) changes from 0 for the standard weekly message to 1 for the daily reminder, the fitted line predicts that learners assigned to daily reminders completed an average of 1.4 more lessons over the two weeks than learners assigned to the standard weekly message. The difference follows from the group means:

$$ 8.5-7.1=1.4\text{ lessons}. $$

Do and conclude (part b). Random assignment makes the two message groups comparable, on average, with respect to other factors at the start of the study. If the assignment and measurement were carried out as described, the results support a causal conclusion about the effect of being assigned daily reminders rather than the standard weekly message for these study participants and this two-week setting. The conclusion concerns assignment to receive the messages; it does not prove that every learner read or followed them, or that every learner completed exactly 1.4 additional lessons.

The participants volunteered rather than being randomly selected from all adult learners. Therefore, random assignment supports a causal comparison for the experiment, but it does not by itself justify generalizing the result to all adult learners. A suitable conclusion stays with the volunteers and study conditions.

Illustrative scoring notes (4 points, not an official rubric): One point for identifying the 0-to-1 change in assigned message condition; one for interpreting the \(1.4\)-lesson slope as a difference in group averages; one for explaining that random assignment supports a causal comparison; and one for keeping the conclusion within the participants and conditions studied, noting that volunteering does not establish broad generalizability. Random selection and random assignment are not interchangeable.

Common Mistakes and Full-Credit Communication

A strong free-response answer is not necessarily long. It is complete because every sentence contributes a relevant fact: what the model predicts, what the study did, or why the design supports—or does not support—a causal claim.

  • Giving a slope without context. “The slope is 0.42” does not identify what changes or in what units. State the predicted change in the response for a one-unit increase in the explanatory variable, naming both variables and their units.
  • Making the slope a guaranteed individual effect. Avoid “each plot will gain” or “a swimmer will improve by.” The fitted line describes a predicted pattern among cases, not a certain outcome for every individual.
  • Calling every regression study an experiment. A regression line can be fitted to observational data or experimental data. Explain whether researchers assigned a treatment or recorded existing values; the equation itself does not answer that question.
  • Stopping at “correlation is not causation.” State the actual design limitation. For observational data, explain that the explanatory variable was not randomly assigned, and give a plausible confounder with a connection to both variables.
  • Offering a confounder with no explanation. Name the variable and say how it could relate to the explanatory variable and to the response. A list of possible factors without those connections does not show why the factor matters.
  • Overstating what random assignment proves. Random assignment can support a causal comparison, but it does not guarantee an exact effect for every individual or make a volunteer sample representative of a broader population.
  • Confusing the treatment with compliance. In an experiment, the assigned condition is what researchers can compare. Unless the study supports a stronger statement, describe the effect of assignment rather than assuming everyone followed or used the treatment.

Before submitting, check that the slope sentence includes the cases, both variables, and the units. Then check that the causation sentence names the study design and explains its consequence. If the data are observational, make the limitation explicit and support it with a plausible confounder. If the study is randomized, connect the causal conclusion to the treatment assigned and keep its scope within the participants and conditions studied.

Key takeaway: A regression free response earns clarity by separating the model’s predicted association from a causal conclusion. Interpret the slope in context, identify whether the explanatory variable was assigned, and explain how that design supports or limits the claim.

Check Your Understanding

For each question, write a concise response that interprets the result and addresses the study design.

  1. A researcher records daily screen time in hours and sleep duration in minutes for 55 volunteers. The fitted slope for predicting sleep duration from screen time is \(-12\). Interpret the slope in context.
  2. In the study in question 1, participants choose their own screen time. Explain why the regression alone does not establish that reducing screen time causes longer sleep. Name a plausible confounder and connect it to both variables.
  3. A coach randomly assigns athletes to one of two warm-up routines and compares their timed sprint results. Explain what feature of the design supports a causal comparison, and name one limit that remains if athletes volunteered for the study.
  4. A fitted line predicts a response difference of 3 points for a 20-minute difference in practice time. Explain why this prediction should not be written as a guaranteed 3-point gain for each person who practices 20 minutes longer.
  5. A student writes, “There is a relationship, so the explanatory variable causes the response.” Identify the reasoning error and describe what study-design detail the student should check.