Tutorials › AP Statistics › Regression Context Mixed Practice Set

Regression and context · Tutorial 940 of 1000

Regression Context Mixed Practice Set

Use a four-part claim audit to interpret regression results and keep conclusions aligned with confounding, study design, and the population actually studied.

Intermediate 9 min read

What You'll Learn

  • Audit a regression claim by checking its interpretation, causal wording, confounding, and population scope.
  • Explain how a plausible confounder could be related to both variables in an observational study.
  • Distinguish what random assignment supports from what random sampling supports.
  • Write slope interpretations with the cases, variables, units, and appropriate associational or causal wording.
  • Revise overgeneralized or overconfident claims into conclusions supported by the study design.

Audit the Whole Claim, Not Just the Regression Line

A regression question may give you a slope and ask what it means. Another may ask whether a news claim is justified, whether a third variable could explain the pattern, or whether the result applies beyond the people or places studied. In mixed practice, these tasks belong together: a mathematically correct interpretation can still overstate causation or generalizability.

In “Free-Response Practice on Regression in Context,” we organized an answer around the model and the study design. This tutorial adds a broader check: audit the interpretation, causal wording, possible confounding, and population scope separately. These are related questions, but evidence for one does not automatically answer the others.

Key strategy: Before accepting a regression claim, ask four questions: What does the model predict, and for which cases? Does the wording claim association or causation? Could a third variable help explain the association? What population, if any, does the sampling method support generalizing to?

Keep two design features distinct. Random assignment can support a causal comparison because it helps create comparable treatment groups. Random selection from a defined population can support generalizing results to that population. A study can have either feature, both, or neither. One does not substitute for the other.

A useful way to organize a response is to write one sentence for each part of the audit. Start with the cases and the regression interpretation. Then describe the study design and the causal limit or support. If the study is observational, identify a plausible confounder and explain its connections. Finish by naming the group to which the result reasonably applies. This order helps keep a strong result in one area from being used as evidence in another.

Worked Mixed-Practice Scenarios

Worked Example: Solar Capacity and Monthly Electricity Costs

Fictional AP-style question. A neighborhood energy group asks 64 homeowners who volunteered for a survey about their solar installations and electricity bills. For each home, it records solar-panel capacity in kilowatts and the average monthly electricity cost in dollars. No panel capacity is assigned. The fitted line for predicting monthly cost from panel capacity is \(\hat{y}=186-14.5x\). A report says, “Each additional kilowatt of solar capacity saves homeowners $14.50 per month, and the finding applies to homeowners throughout the state.”

State and plan. The cases are the 64 volunteer homeowners. The explanatory variable is solar-panel capacity, in kilowatts, and the response is average monthly electricity cost, in dollars. I will interpret the slope as a prediction from the fitted line, then check whether the data support the report’s causal wording and broad population claim.

Do: interpret and audit the wording. The slope is \(-14.5\) dollars per kilowatt. A careful interpretation is: Among the 64 volunteer homeowners studied, the fitted line predicts an average monthly electricity cost $14.50 lower for each additional kilowatt of solar-panel capacity. The negative sign indicates a predicted decrease in cost as capacity increases. The sentence describes an association in the observed homes; by itself, it does not say that increasing a home’s capacity will cause its bill to fall by that amount.

The report changes an associational prediction into a causal claim with the verb “saves.” But this was an observational study: homeowners already had different panel capacities, and researchers recorded them. A plausible confounder is household electricity use. Larger households may use more electricity and have different choices about panel capacity; their electricity use also affects the monthly bill. Thus household use could be related to both the explanatory variable and the response, mixing its influence into the observed relationship.

Conclude: assess generalization. The homeowners volunteered; they were not randomly selected from all homeowners in the state. Their survey could differ from the experiences of homeowners who did not volunteer, and the study gives no basis for claiming that the result applies statewide. A careful conclusion is: For the volunteer homeowners surveyed, greater solar-panel capacity was associated with lower average monthly electricity costs, with the fitted line predicting a $14.50 lower cost per additional kilowatt. Because capacity was not randomly assigned, the study does not establish that adding capacity causes this decrease. Because the homeowners volunteered, the result should not be generalized to all state homeowners.

Illustrative scoring notes (4 points, not an official rubric): One point for a slope interpretation with the cases, variables, and units; one for identifying the recorded, observational design and avoiding a causal conclusion; one for explaining a plausible confounder’s relation to both variables; and one for recognizing that volunteers do not justify a statewide generalization. Merely saying “there could be bias” does not explain the confounding or sampling limitation.

Worked Example: Compost Reminders and Household Composting

Fictional AP-style question. A city recruits 90 households that have already signed up for its composting information program. Researchers randomly assign 45 households to receive weekly text reminders and 45 to receive the standard monthly newsletter. After eight weeks, they record the amount of food waste each household composted, in kilograms. Let \(x=0\) represent the newsletter and \(x=1\) represent weekly texts. The group means are 5.8 kilograms for \(x=0\) and 7.1 kilograms for \(x=1\), and the fitted line is \(\hat{y}=5.8+1.3x\). A city announcement says, “Weekly texts cause every household to compost 1.3 more kilograms, so the same increase can be expected across the city.”

State and plan. The cases are the 90 participating households. The explanatory variable is assigned message type, coded 0 or 1, and the response is kilograms of food waste composted in eight weeks. Because researchers randomly assigned the messages, this is an experiment. I will interpret the slope as a difference between assigned groups, then consider what random assignment and the volunteer recruitment do—and do not—support.

Do: interpret and assess causation. For a change from \(x=0\) to \(x=1\), the fitted line predicts a response difference of

$$ (5.8+1.3)-5.8=7.1-5.8=1.3\text{ kilograms}. $$

A complete interpretation is: Among the participating households, those randomly assigned to weekly text reminders composted an average of 1.3 kilograms more food waste over eight weeks than those assigned to the monthly newsletter. Random assignment supports a causal comparison of assignment to the two message conditions for these participants, assuming the study was carried out as described. The conclusion concerns the effect of being assigned weekly reminders, not a guaranteed increase for each household. Individual households can differ in their responses.

Conclude: limit the population claim. The households had signed up for the city’s program; they were not randomly selected from all city households. Random assignment helps support a causal conclusion within the experiment, but it does not make the participants representative of the city. The announcement’s claim that the increase can be expected across the city is therefore too broad. Also, “every household” incorrectly turns a group-average difference into a certain individual outcome.

A more defensible statement is: For the participating households in this eight-week experiment, assignment to weekly text reminders resulted in an average of 1.3 more kilograms of food waste composted than assignment to the monthly newsletter. The experiment supports a causal comparison for these participants, but the volunteer group does not establish that the same average effect applies across the city or to every household.

Illustrative scoring notes (4 points, not an official rubric): One point for interpreting the 1.3-kilogram slope as a difference between assigned groups; one for connecting random assignment to a causal comparison; one for rejecting an individual guarantee; and one for limiting generalization because participants signed up rather than being randomly selected from city households. A response that says simply “it was randomized, so it applies to everyone” confuses assignment with selection.

Worked Example: Shade and Water Temperature at Public Swimming Areas

Fictional AP-style question. A county has 40 public swimming areas. Staff randomly select 24 of them and, during the same week in July, record the percentage of each area’s shoreline shaded at midday and the water temperature in degrees Celsius. Shade is not assigned. The fitted line for predicting temperature from shade percentage is \(\hat{y}=27.8-0.07x\). A community newsletter writes, “Adding 10 percentage points of shade will lower water temperature by 0.7°C at swimming areas everywhere.”

State and plan. The cases are the 24 randomly selected swimming areas. The explanatory variable is shoreline shade percentage, and the response is water temperature in degrees Celsius. I will translate the slope into the requested 10-percentage-point difference, then distinguish what random selection supports from what the observational design can show about cause.

Do: interpret the slope and wording. The slope is \(-0.07\)°C per percentage point of shade. For a 10-percentage-point difference, the fitted line predicts

$$ -0.07(10)=-0.7^\circ\text{C}. $$

Thus, among the sampled swimming areas during the week studied, an area with 10 percentage points more shoreline shade is predicted by the fitted line to have water temperature 0.7°C lower, on average. This is a predicted comparison across areas, not a guarantee about changing shade at one area. Staff observed existing shade rather than assigning it, so the data do not establish that adding shade causes a 0.7°C decrease.

A plausible confounder is water depth. Deeper areas may have different shoreline shade and may also have different water temperatures. If depth is related to both shade and temperature, it could contribute to the observed association. This is a possible explanation to consider, not proof that depth caused the pattern.

Conclude: identify the supported scope. Random selection from the county’s 40 public swimming areas supports generalizing an association to that defined group, under the conditions represented by the sampling and measurement process. It does not support a causal conclusion, because shade was not assigned. Nor does it support “everywhere”: the sampled frame consists of 40 county swimming areas observed during one July week, not swimming areas everywhere or all seasons.

A careful replacement is: For the county’s 40 public swimming areas during the July week studied, the data provide evidence of an association between greater shoreline shade and lower water temperature. The fitted line predicts a 0.7°C lower temperature for a 10-percentage-point increase in shade. Because shade was observed rather than assigned, this association does not show that adding shade causes the decrease.

Illustrative scoring notes (4 points, not an official rubric): One point for translating the slope into a contextual 10-percentage-point prediction; one for identifying that observed shade does not establish causation; one for naming and connecting a plausible confounder; and one for using the random selection to describe the limited population scope while rejecting “everywhere.” Random selection supports generalization to the sampling frame, not automatically to other places or times.

Common Mistakes and a Claim-Audit Checklist

These scenarios show why a regression answer should be checked along more than one dimension. A study can support a causal comparison without supporting broad generalization. It can support generalizing an association without establishing causation. A strong interpretation does not repair a weak sampling method, and a strong sampling method does not turn an observational relationship into a causal effect.

  • Using causal verbs for recorded variables. “Causes,” “increases,” “saves,” and “leads to” make causal claims. For an observational study, describe what values tend to occur together or what the fitted line predicts. Explain that the explanatory variable was not randomly assigned.
  • Naming a confounder without explaining it. A full explanation identifies a specific variable and says how it could be related to both the explanatory variable and the response. “Other factors could matter” does not demonstrate why the association might be confounded.
  • Confusing a difference between cases with a change within one case. A slope describes a predicted pattern across observations. Do not turn “areas differing by 10 percentage points” into a guaranteed temperature change at a particular area.
  • Claiming that random assignment makes a sample representative. Assignment concerns how participants are placed into conditions; selection concerns how they enter the study. A volunteer experiment may support a causal comparison among its participants without supporting generalization to a wider population.
  • Generalizing beyond the sampling frame. Name the group from which cases were selected. If sites were randomly selected from county swimming areas, do not silently extend the conclusion to every swimming area, all seasons, or a different region.
  • Giving a technically correct slope without its setting. State what the cases are, name the explanatory and response variables, include their units, and say whether the line predicts a response difference. A bare number such as “minus 0.07” is not a contextual interpretation.

Before finishing a mixed-practice response, make a quick four-part audit: interpretation—did you name the cases, variables, units, and predicted change? Wording—does the sentence claim association or cause, and does that match the design? Confounding—for observational data, did you explain how a plausible third variable relates to both variables? Scope—does the sampling method support the population named in the conclusion?

Key takeaway: Interpret the fitted line in context, then evaluate causal wording, plausible confounding, and population scope as separate parts of the claim. Random assignment can support causation; random selection can support generalization. Neither should be claimed unless the study design provides it.

Check Your Understanding

For each scenario, write a response that interprets the regression result and checks the claim against the study design.

  1. A researcher surveys volunteer gardeners and finds a slope of \(2.4\) for predicting weekly vegetable harvest in kilograms from weekly watering time in hours. Interpret the slope without making a causal claim.
  2. In question 1, name a plausible confounder and explain how it could be related to both watering time and harvest.
  3. A school randomly assigns volunteer students to one of two study-planner formats. Explain what random assignment supports and why volunteering still matters when considering generalization.
  4. Staff randomly select 18 trails from a county’s 30 trails and record trail width and the number of visitors. Explain what random selection may support and why the observed regression does not show that widening a trail causes more visitors.
  5. A student writes, “Our randomly assigned experiment proves the treatment works for all residents.” Identify the two distinct claims in this sentence and state what additional design information is needed to assess each one.