Score the Evidence in the Response
A regression free-response answer may contain the right numbers but still miss credit because it does not explain what those numbers mean. It may also earn credit for one correct statement while making a separate, unsupported claim. Scoring carefully means checking each requested idea on its own—not deciding that an answer “sounds good” overall.
In “Common Communication Errors in Regression,” you practiced identifying wording that misstates a statistic or overreaches. Here, you will apply a scoring rubric to sample responses. The key new technique is to translate each rubric line into a question you can answer using the student’s actual words: Is the required idea present? Is it correct? Does it include the context, units, or limit the criterion requires?
The rubrics in this tutorial are teaching examples, not official College Board scoring rubrics. Actual AP questions have their own rubrics, and the wording of those rubrics determines what earns credit. Do not assume that every regression question has the same number of points or requires the same statistics.
Turn the Rubric Into a Checklist
Before scoring a response, identify what the question asks for and split the rubric into separate criteria. For example, “interpret the slope” is not the same task as “interpret \(r^2\).” A correct slope statement cannot make up for an incorrect \(r^2\) statement, and a correct statistic does not automatically earn credit for a missing interpretation.
Use the criterion itself to set the standard. If it asks for predicted change in the response per unit of the explanatory variable, look for both variables, direction, and units. If it asks for the proportion of variation explained, look for variation in the response, the cases represented, and the linear model. Avoid adding requirements that the rubric does not include, but do not award a point for an idea that is only implied or nearly correct.
List the distinct ideas the question requires. Note whether it asks for a calculation, interpretation, comparison, or limitation.
Find the exact words, numbers, or calculations in the response that address that criterion.
Ask whether the evidence is statistically correct and includes the context or units the criterion calls for.
Award credit for each satisfied criterion; do not let a correct statement cancel an error elsewhere.
For missed credit, identify what is absent or incorrect and state the smallest change that would make the idea complete.
A practical scoring table can keep this process consistent. “The student understands the topic” is not evidence for a point. A quoted phrase or shown calculation is. If a rubric awards one point only when several pieces are present, apply that rule as written rather than awarding a point for just one piece.
Worked Example: Score a Set of Interpretations
Worked Example: Light and Plant Growth
Situation. In an invented project, researchers recorded daily light exposure and plant growth for 28 greenhouse plants. Let \(x\) be hours of light per day and \(y\) be weekly growth in centimeters. The observed light exposure ranged from 2 to 9 hours. The regression output is
A question asks for interpretations of the slope, \(r^2\), and \(s\). For practice, use this three-point rubric: one point for a contextual slope interpretation with units and predicted direction; one point for interpreting \(r^2\) as the proportion of variation in growth accounted for by the linear model; and one point for describing \(s\) as typical residual size in centimeters without making it a guarantee.
A student responds: “Each extra hour of light increases growth by 0.62. The model accounts for 65.61% of growth variation. The plants are all within 1.1 cm of the line.”
Score the slope criterion. The response has the numerical slope and mentions an extra hour, but it omits the response units, centimeters per hour. “Increases” may also imply a guaranteed change rather than a predicted change from the fitted line. A complete statement would be: “For these 28 plants, each additional hour of daily light is associated with a predicted increase of 0.62 centimeters in weekly growth, according to the fitted line.” Under the stated criterion, the original does not earn the point because its interpretation is incomplete.
Score the \(r^2\) criterion. The calculation is consistent: \(0.81^2=0.6561\), and \(0.6561\times100\%=65.61\%\). The student names variation in growth and says the model accounts for that variation. This earns the point under the stated criterion. A more fully contextual version would specify that it is variation among the 28 plants and that the model is the linear regression of weekly growth on daily light exposure.
Score the \(s\) criterion. “The plants are all within 1.1 cm of the line” is not a valid interpretation. It turns a typical residual size into a guarantee about every plant. This criterion earns no point. A full-credit version is: “For these plants, residuals typically have a size of about 1.1 centimeters; this does not guarantee that every plant’s growth is within 1.1 centimeters of its fitted value.”
Conclude. The response earns 1 of the 3 available points under this practice rubric: credit for \(r^2\), but not for the incomplete slope interpretation or the incorrect guarantee about \(s\). The score reflects the stated criteria, not a general judgment about whether the student “knows regression.”
Notice that the \(r^2\) statement is brief yet earns credit because it correctly states the requested meaning. You do not need to penalize every omission in the response if that omission is not part of the criterion being scored. At the same time, a number alone is not an interpretation when the rubric explicitly asks what the number means.
Worked Example: Score a Prediction and Its Scope
Worked Example: Delivery Time and Distance
Situation. In an invented set of 32 local deliveries, let \(x\) be distance in kilometers and \(y\) be delivery time in minutes. The observed distances ranged from 3 to 18 kilometers. A fitted line is \(\hat{y}=7.5+2.4x\), with \(s=4.0\) minutes. A question asks for the predicted delivery time at 12 kilometers and whether the prediction uses extrapolation.
Use this four-point practice rubric: one point for substituting 12 into the fitted line and finding the predicted time; one point for interpreting that value as a predicted time in context; one point for comparing 12 with the observed \(x\)-range and correctly classifying the prediction; and one point for describing the role of \(s\) accurately if discussing prediction accuracy.
A student writes: “At 12 km, delivery time is \(7.5+2.4(12)=36.3\), so the delivery takes exactly 36.3 minutes. Since 12 is between 3 and 18, this is not extrapolation. The model’s error is 4 minutes.”
Score the calculation and prediction criteria. The substitution is correct:
The calculation earns its point. The contextual interpretation does not earn its point as written: “takes exactly” presents a fitted prediction as a guaranteed observed time. A full-credit interpretation would say, “For a delivery distance of 12 kilometers, the fitted line predicts a delivery time of 36.3 minutes.”
Score the range comparison. The student correctly compares 12 kilometers with the observed range of 3 to 18 kilometers and identifies the prediction as not extrapolation. This earns the point. As emphasized in “Common Errors About Extrapolation,” the classification depends on the requested explanatory-variable value, not on how large or precise the predicted response seems.
Score the \(s\) statement. “The model’s error is 4 minutes” is too definite and does not explain what \(s\) describes. This criterion does not earn its point. The output value \(s=4.0\) minutes means that residuals typically have a size of about 4 minutes for these deliveries. It does not give the error for this particular delivery or guarantee that its delivery time is within 4 minutes of 36.3.
Conclude. The response earns 2 of 4 points: one for the fitted-value calculation and one for the observed-range comparison. Replacing “takes exactly” with “the fitted line predicts” would earn the interpretation point, while describing \(s\) as typical residual size rather than “the model’s error” would earn the final point.
This example shows why one sentence may contain both creditworthy and incorrect parts. The correct range check does not repair the overstatement that the delivery takes exactly 36.3 minutes. Score each requested idea on its own evidence.
Worked Example: Find the Errors Before Awarding Credit
Worked Example: Screen Brightness and Battery Use
Situation. In an invented observational analysis of 24 tablets used in one school, let \(x\) be screen brightness in percentage points and \(y\) be battery charge used per hour, also in percentage points. Brightness values ranged from 25 to 85. The fitted line is \(\hat{y}=0.7+0.09x\), with \(r=0.76\) and \(r^2=0.5776\). A question asks for interpretations of the slope and \(r^2\), and asks whether the data show that changing brightness causes battery use to rise.
Use this five-point practice rubric: one point for interpreting the slope as predicted change in battery use per brightness percentage point, with units; one for describing the direction and strength of the linear association using \(r\); one for correctly interpreting \(r^2\) as a proportion of variation in the response accounted for by the model; one for recognizing that these observational data do not establish causation; and one for keeping the conclusion within the studied cases and setting.
A student writes: “Increasing brightness by 1% causes battery use to rise by 0.09%. The correlation says the model is 76% accurate. Since \(r^2=57.76\%\), brightness explains that much battery use for all tablets. These numbers prove the same effect will happen in any school.”
Score the slope criterion. The number \(0.09\) is the slope, but the response incorrectly makes a causal claim and does not clearly state the response change per one percentage point of brightness. The criterion earns no point. A stronger interpretation is: “Among the tablets studied, each additional percentage point of screen brightness is associated with a predicted increase of 0.09 percentage points in battery charge used per hour, according to the fitted line.”
Score the \(r\) criterion. The response uses \(r=0.76\) to claim that the model is “76% accurate.” That is not what correlation measures. The positive \(r\) indicates a positive linear association, and its magnitude describes the strength of that linear association. The statement does not earn the point. A complete interpretation would refer to the positive linear association between screen brightness and battery charge used per hour among these tablets.
Score the \(r^2\) criterion. The calculation is \(0.76^2=0.5776\), or \(57.76\%\). But “explains that much battery use” is not a precise interpretation of response variation, and the claim about all tablets goes beyond the cases studied. As written, the criterion earns no point. A full-credit version is: “About 57.76% of the variation in battery charge used per hour among these 24 tablets is accounted for by the linear regression of battery use on screen brightness.”
Score causation and scope. The student says the results “prove” a causal effect and generalize to any school. Both claims are unsupported. Brightness was observed, not assigned, so the analysis establishes an association rather than cause and effect. The 24 tablets came from one school, so the results do not automatically apply to tablets in every school. Neither criterion earns a point.
Conclude. The response earns 0 of 5 points under this practice rubric. Several relevant numbers are present, but none is used in the required way. A full-credit response would distinguish slope from \(r\) and \(r^2\), describe the observed association without claiming causation, and restrict the conclusion to the tablets and setting represented.
Common Scoring Mistakes and AP Exam Tips
- Scoring the intention instead of the words. Do not award a point because the student may have meant the right thing. If the criterion requires response units and the answer gives none, check whether the rubric permits credit without them.
- Letting one correct number carry several criteria. A correct \(r^2\) value does not automatically earn the interpretation point. A correct fitted value does not automatically earn a point for checking extrapolation.
- Using an all-or-nothing impression. A response can earn some points and miss others. Mark each criterion separately, especially when a sentence combines correct evidence with an incorrect conclusion.
- Demanding extra wording not in the rubric. Apply the rubric as written. Do not withhold credit for a stylistic preference if the required statistical meaning is clearly present.
- Confusing an error with an omission. A missing unit is an omission; calling \(r\) an accuracy percentage is an incorrect interpretation. Name the specific issue when explaining lost credit.
- Giving credit for a claim that overreaches. “Causes,” “always,” and “for every case” can change the meaning of an otherwise relevant result. Check whether the data design and the rubric support those claims.
When you review your own response, underline the words that satisfy each rubric criterion. If you cannot point to the evidence, revise the answer. A full-credit response is not necessarily long: it is complete for the task, statistically accurate, and clear about context, units, and limits where required.
Check Your Understanding
For each situation, decide which rubric criterion is met or missed, and state what wording would make the answer more complete.
- A rubric asks for the slope in context. A student writes, “The slope is 3.2.” What information is still needed for a contextual interpretation?
- A regression has \(r=-0.68\). A response says, “The model is 68% accurate.” Which criterion would this miss, and what does \(r\) describe instead?
- A student correctly calculates a fitted value but says the observed response “will equal” that value. Should the calculation and interpretation necessarily receive the same score? Explain.
- A response interprets \(r^2\) correctly but claims that the relationship applies to every person in the country, although the data came from one school. Which part may earn credit, and which claim needs revision?
- Why should a scorer use the question’s actual rubric instead of assuming every regression free response awards points for the same four statistics?