Tutorials › AP Statistics › Variables, Units, and Meaning in a Regression Model

Regression and context · Tutorial 921 of 1000

Variables, Units, and Meaning in a Regression Model

Learn to read a regression setting carefully by naming who or what each observation describes, what is measured, and the units and roles of the variables.

Intermediate 10 min read

What You'll Learn

  • Identify the individuals represented by the rows of a regression data set.
  • Distinguish the response variable from the explanatory variable.
  • State the units of each variable, including rates and percentages.
  • Label \(x\) and \(y\) consistently when reading a regression equation.
  • Recognize when changing variable roles changes the question being modeled.
  • Avoid treating an explanatory variable as proof of a cause.

Start With the Observations and the Question

A regression setting describes paired measurements: for each individual or case, there is a value of an explanatory variable and a corresponding value of a response variable. Before interpreting a regression result such as \(r^2\), make sure you know what one case represents and what each variable measures. Those details determine what a statement about the model actually means.

In the previous tutorial, “\(r^2\) Mixed Practice Set,” interpretations named the response, the explanatory variable, and the observed cases. This tutorial focuses on identifying those parts from a description. The key is to distinguish who or what is observed from what is measured, and then to track the units of each measurement.

Definition: An individual (also called a case or observation) is the person, object, place, or other unit described by one row of data. A variable is a characteristic measured for each individual. In a regression setting, the explanatory variable is used to explain or predict the response, and the response variable is the outcome being explained or predicted.

A description may contain several nouns and measurements. Do not assume that the most prominent noun is a variable: “plants,” “visits,” and “parks” might name the individuals, while sunlight, waiting time, or attendance counts are variables measured for them. Similarly, a setting may mention extra facts that are not part of the regression model.

Use the study’s question to identify variable roles. If a study asks whether sunlight amount helps predict plant growth, sunlight is explanatory and growth is the response. If the question instead asks whether growth helps predict later sunlight exposure, the roles would be reversed. The measurements themselves do not carry permanent labels; their roles depend on what the model is intended to describe.

Identification routine: For a regression setting, ask: (1) What does one data row represent? (2) What two quantities are recorded for each row? (3) Which quantity is used to explain or predict the other? (4) What are the units of each quantity?

The units should describe the measurement as it is recorded, not just its general topic. “Time” is incomplete if the data record minutes. “Visitors” is incomplete if the recorded variable is visitors per Saturday. If a variable is a percentage, state what percentage is being measured; where helpful, distinguish percentage from percentage points. Clear units are essential when interpreting predictions, residuals, or the slope of a fitted line.

As in earlier work with the regression equation \(\hat{y}=a+bx\), \(x\) represents the explanatory variable and \(y\) the response variable; \(\hat{y}\) is a predicted response. When writing an equation in context, label the variables and their units so a reader can tell what values are being predicted. The slope’s units are response units per explanatory-variable unit. The intercept has response units. These unit labels do not, by themselves, show that changing \(x\) causes a change in \(y\).

Worked Examples: Name the Cases, Variables, and Units

Worked Example: Sunlight and Tomato-Plant Growth

A fictional gardening project records data for 36 tomato plants. For each plant, the team records the average number of hours of direct sunlight it receives per day during a six-week growing period and the plant’s dry mass, in grams, at the end of the period. The team wants to use sunlight to predict dry mass. Identify the individuals, the variables, their roles, and their units.

Identify the individuals. One individual is one tomato plant. Each of the 36 rows describes a different plant, not a day of sunlight or a gram of plant material.

Identify the variables. The two recorded variables are average daily direct-sunlight exposure and end-of-period dry mass. Sunlight exposure is measured in hours per day. Dry mass is measured in grams per plant.

Assign the roles. Because the project uses sunlight to predict dry mass, average daily sunlight is the explanatory variable, \(x\), and dry mass is the response variable, \(y\). The data pair for each plant consists of its sunlight value and its dry-mass value.

A context-labeled model could be written as \(\widehat{\text{dry mass}}=a+b(\text{sunlight})\). Here, \(x\) is sunlight in hours per day, and \(\hat{y}\) is predicted dry mass in grams per plant. If a fitted slope were reported, its units would be grams per plant per additional hour of average daily sunlight. The unit description tells us how to read the numerical slope; it does not establish a causal effect.

Conclusion. The 36 tomato plants are the individuals. Average daily direct sunlight, in hours per day, is the explanatory variable; end-of-period dry mass, in grams per plant, is the response variable. The project’s prediction direction comes from its stated question.

Worked Example: Queue Length and Waiting Time

A fictional transit agency records 42 passenger visits at a service desk. At the moment each passenger joins the line, an observer records the number of people already waiting. The observer also records the passenger’s waiting time from joining the line until reaching the desk, in minutes. The agency wants to predict waiting time from the number of people already waiting. Identify the cases and variables precisely.

Identify the individuals. One individual is one passenger visit to the service desk. The recorded cases are the 42 visits; the individuals are not the transit agency, the desk, or the people who happen to be waiting in line.

Identify the variables and units. The explanatory variable is the number of people already waiting when that passenger joins the line. Its unit is people waiting, recorded as a count. The response variable is that passenger’s waiting time, measured in minutes.

Check the pairing. Each row must keep together the queue count and waiting time for the same visit. For example, the count recorded when one passenger joins must be paired with that passenger’s waiting time—not with another passenger’s wait.

If \(x\) denotes the number of people already waiting and \(y\) denotes the waiting time, a fitted equation predicts \(\hat{y}\), a waiting time in minutes, for a specified queue count. The equation does not predict the number of people in line from a waiting time; that would make a different quantity the response.

Conclusion. The individuals are the 42 passenger visits. Queue length, measured as people already waiting, is explanatory; waiting time, measured in minutes, is the response. Both the moment of recording the queue count and the person whose wait is measured matter for describing the variables accurately.

Worked Example: Park Features and Saturday Attendance

A fictional community group studies 30 parks. For each park, volunteers record the number of playground structures and the number of visitors observed on one specified Saturday. They want to use the playground-structure count to predict the visitor count. State the individuals, the variables, the roles, and the units. Also explain what one row’s visitor measurement describes.

Identify the individuals. One individual is one park. The 30 parks are the observed cases. The volunteers are not the individuals in this data set, and individual visitors are not the cases either.

Identify the variables and units. The number of playground structures is a count, measured in structures per park. The response is the number of visitors observed at that park on the specified Saturday, measured in visitors per park for that Saturday.

Assign the roles. The stated prediction goal makes playground-structure count the explanatory variable and Saturday visitor count the response. For each park, its structure count must be paired with its visitor count from that same park and the same observation day.

The time wording is part of the response’s meaning. The outcome is not a park’s annual attendance or its typical daily attendance; it is the count observed on one specified Saturday. If the study had counted visitors on a different day or over a whole year, it would be measuring a different response variable.

Conclusion. The parks are the individuals. Playground structures per park is the explanatory variable, and visitors per park on the specified Saturday is the response. Saying simply “visitors” would omit the observation period and leave the measurement less clear.

Common Mistakes and Full-Credit Wording

A careful answer does not need to be long, but it should make the case, variable roles, and units unmistakable. A useful check is whether another reader could tell what one row means and what a value of each variable represents.

  • Confusing the individual with a variable. In the park example, a park is an individual; visitor count is a variable measured for each park. Full-credit wording explicitly says what one case represents.
  • Giving a topic instead of a variable. “Growth” or “time” may be too vague. State the recorded quantity, such as dry mass in grams or waiting time in minutes.
  • Leaving out the unit or observation period. “Visitors” does not fully describe visitors observed on one Saturday. State the relevant unit and time period when they are part of the measurement.
  • Reversing the response and explanatory roles. The response is the quantity the model is intended to explain or predict. Use the question’s wording to decide which variable plays that role.
  • Pairing values from different cases. Each explanatory value must be matched with its corresponding response value for the same individual. Otherwise, the row no longer describes one case.
  • Assuming the explanatory variable causes the response. “Explanatory” names its role in the model, not proof of cause and effect. The project’s design and evidence determine what causal claims, if any, are justified.
  • Using role labels without context. Saying only “\(x\) is explanatory and \(y\) is response” is less informative than naming each quantity and its units.

For example, a complete identification could read: “Each row represents one tomato plant. Average daily direct sunlight, in hours per day, is the explanatory variable, and end-of-period dry mass, in grams per plant, is the response.” This sentence tells the reader who or what is observed, what is measured, how the model is oriented, and what units the measurements use.

If a regression equation or computer output is provided, keep the labels aligned with the context. Confirm that \(x\) is the explanatory quantity named in the question and that predicted \(\hat{y}\) has the response’s units. This is also a useful check on interpretation: a predicted value should have units that make sense for the response, not the explanatory variable.

Key takeaway: Begin a regression interpretation by naming the individual represented by one observation, then identify the response and explanatory variables with their units. The study’s question determines their roles, and the paired values must belong to the same case.

Check Your Understanding

For each setting, name one individual, both variables with units, and which variable is explanatory and which is the response.

  1. A fictional study records the age, in years, and monthly electricity use, in kilowatt-hours, for 48 apartment refrigerators. The goal is to predict electricity use from age.
  2. A school project records the number of minutes students spend commuting and the number of days they arrive on time during a 20-day period. The goal is to predict on-time days from commute time. What is one individual, and what unit or time period must be included for the response?
  3. A fictional research team records a park’s area in hectares and the number of bird species observed there. It uses area to predict species count. Identify the individuals and the variables’ roles.
  4. In a study of 25 delivery routes, a team records route distance in kilometers and fuel used in liters. If the question changes from predicting fuel use from distance to predicting distance from fuel use, what changes and what stays the same?
  5. A student says, “The explanatory variable causes the response because it is used to explain it.” Explain what is wrong with this claim.