Tutorials › AP Statistics › Correlation Mixed Practice Set

Correlation · Tutorial 840 of 1000

Correlation Mixed Practice Set

Work through a mixed set of correlation problems that combines calculation, contextual interpretation, and careful critique of what \(r\) can show.

Intermediate 10 min read

What You'll Learn

  • Calculate \(r\) from paired observations using centered deviations.
  • Verify a calculated correlation with a second method or calculator output.
  • Interpret the direction and strength of a correlation in context.
  • Recognize when a near-zero correlation hides a curved pattern.
  • Critique claims that treat correlation as a cause or a guaranteed outcome.

A Mixed Practice Routine for Correlation

In “Writing a Correlation Conclusion on the AP Exam,” you practiced turning a correlation into a complete, contextual statement. This tutorial combines that skill with calculating \(r\) and deciding whether it is an appropriate summary of the data. The aim is not just to get a number: it is to connect the number to the paired observations, the scatterplot, and the limits of the study.

For each scenario, use the same routine: identify the two quantitative variables and the individuals; inspect the pattern’s form and unusual features; calculate or check \(r\); then interpret the result in context. As in “Correlation Measures Only Linear Association,” remember that \(r\) summarizes linear association. A calculation can be correct while a conclusion based on it is incomplete or misleading.

Mixed-practice checklist: Keep each \(x\)-value paired with its matching \(y\)-value. Check whether a linear summary makes sense for the pattern. Calculate \(r\) carefully, verify it, and interpret its direction and strength in context without claiming causation or certainty for individual cases.

Calculate, Verify, and Interpret

For a small data set, one way to calculate the correlation is to compare each observation with the sample mean of its variable. Multiply the centered \(x\)- and \(y\)-values for each individual, add those products, and divide by the square root of the product of the two sums of squared deviations.

$$ r=\frac{\sum (x_i-\bar{x})(y_i-\bar{y})} {\sqrt{\sum (x_i-\bar{x})^2\sum (y_i-\bar{y})^2}} $$

This formula gives one number for the paired observations. The numerator reflects whether above-average \(x\)-values tend to pair with above-average or below-average \(y\)-values. The denominator scales the result so that, when it is defined, \(r\) is between \(-1\) and \(1\). In “Finding \(r\) With LinReg on the Calculator,” you learned to read \(r\) directly from the calculator’s LinReg output. That provides a useful independent check on hand arithmetic.

A useful verification is to calculate \(r\) from the centered-deviation formula and then check the result with LinReg using the same paired lists. Confirm that the calculator’s \(r\) has the same sign and a matching value after rounding. If the answers disagree, first check that the values are correctly paired and that the calculator lists do not contain extra or missing entries.

Worked Examples: Calculation, Interpretation, and Critique

Worked Example: Study Time and a Practice-Quiz Score

A fictional data set records the number of hours five students spent studying for a practice quiz and each student’s score, in points. The paired observations are \((1,52)\), \((2,55)\), \((3,59)\), \((4,61)\), and \((5,68)\). Calculate and interpret \(r\).

The individuals are the five students. Let \(x\) be study time in hours and \(y\) be quiz score in points. The means are \(\bar{x}=3\) hours and \(\bar{y}=59\) points. The centered values and their products are:

Study time \(x\)Score \(y\)\(x-\bar{x}\)\(y-\bar{y}\)Product
152-2-714
255-1-44
359000
461122
5682918

The products sum to \(38\). The squared \(x\)-deviations sum to \(4+1+0+1+4=10\), and the squared \(y\)-deviations sum to \(49+16+0+4+81=150\). Substitution gives:

$$ r=\frac{38}{\sqrt{(10)(150)}}= \frac{38}{\sqrt{1500}}\approx 0.9812 $$

As a check, LinReg on the same five paired values gives \(r\approx 0.9812\). The squared value is about \(0.9627\), and its positive square root is about \(0.9812\), which also confirms the magnitude and sign. The calculation indicates a very strong positive linear association in these observations. A complete contextual interpretation is: “Among these five students, study time, in hours, and practice-quiz score, in points, have a very strong positive linear association; students who studied more hours tended to earn higher scores.”

This is a tiny fictional data set, so the statement describes these observations; it does not establish a pattern for all students. And because the example records study time rather than assigning it, the association alone does not show that studying more caused the higher scores. A scatterplot should still be checked for unusual features before relying on \(r\).

Worked Example: Screen Time and Sleep Duration

A fictional survey records evening screen time, in hours, and sleep duration, in hours, for five teenagers. The paired values are \((2,30)\), \((4,27)\), \((6,26)\), \((8,20)\), and \((10,17)\). Calculate \(r\), then critique the statement, “Each extra hour of screen time causes teenagers to lose sleep.”

The individuals are the five surveyed teenagers. Here, \(\bar{x}=6\) hours and \(\bar{y}=24\) hours. The centered values are \(-4,-2,0,2,4\) for screen time and \(6,3,2,-4,-7\) for sleep duration. Their paired products sum to \((-4)(6)+(-2)(3)+(0)(2)+(2)(-4)+(4)(-7)=-66\). The squared deviations sum to \(40\) for \(x\) and \(114\) for \(y\). Thus:

$$ r=\frac{-66}{\sqrt{(40)(114)}}= \frac{-66}{\sqrt{4560}}\approx -0.9774 $$

A calculator check using LinReg gives \(r\approx -0.9774\). As a second arithmetic check, \(r^2=4356/4560\approx 0.9553\), so taking the negative square root gives approximately \(-0.9774\). The negative sign matches the pattern: greater recorded screen time tends to go with shorter sleep duration. The data show a very strong negative linear association among these five teenagers.

The student’s claim goes beyond what \(r\) establishes. “Causes” makes a causal claim, and “each extra hour” suggests a fixed effect for every teenager. Neither follows from a correlation. A more careful interpretation is: “Among these five surveyed teenagers, evening screen time and sleep duration have a very strong negative linear association: those with more evening screen time tended to report fewer hours of sleep. Because these are observational data, the association alone does not show that screen time caused shorter sleep.”

The response does not assume that the same association applies to all teenagers. It describes the surveyed individuals and the measured variables. This is the kind of distinction emphasized in “Why Correlation Does Not Imply Causation.”

Worked Example: Temperature Departure and Greenhouse Complaints

A fictional log records, for five days, how far the afternoon temperature was from a comfortable reference temperature and how many visitors complained that the greenhouse felt uncomfortable. Let \(x\) be the signed temperature departure, in degrees Celsius: negative values are cooler than the reference and positive values are warmer. The paired data are \((-2,4)\), \((-1,1)\), \((0,0)\), \((1,1)\), and \((2,4)\). Calculate \(r\) and evaluate the claim, “Temperature departure is unrelated to complaints because \(r=0\).”

The means are \(\bar{x}=0\) degrees and \(\bar{y}=2\) complaints. The centered products sum to \((-2)(2)+(-1)(-1)+(0)(-2)+(1)(-1)+(2)(2)=0\). The squared \(x\)-deviations sum to \(10\), and the squared \(y\)-deviations sum to \(14\). Therefore:

$$ r=\frac{0}{\sqrt{(10)(14)}}=0 $$

LinReg on the same paired values also gives \(r=0\), verifying the calculation. However, the plotted points form a clear U-shaped pattern: complaints are higher for days that are either cooler or warmer than the reference and lower near the reference. There is no linear association in these five observations, but there is a clear curved association.

The claim that temperature departure is “unrelated” is therefore too broad. A suitable conclusion is: “For these five days, there is no linear association between signed temperature departure, in degrees Celsius, and the number of discomfort complaints. However, the pattern is curved, with more complaints at departures in either direction from the reference, so \(r=0\) does not mean the variables have no relationship.”

This example shows why the scatterplot and \(r\) answer related but different questions. The graph reveals the curve; \(r\) summarizes only linear association. As discussed in “Correlation Measures Only Linear Association,” a small or zero correlation cannot rule out a strong nonlinear pattern.

Common Mistakes and AP Exam Tips

  • Breaking the pairs. Each \(x\)-value must stay matched with the \(y\)-value from the same individual. A mismatched list changes the data and may change \(r\). Check the paired observations before calculating.
  • Reporting a calculation without interpreting it. A value such as \(r=-0.98\) needs a contextual statement. Name the individuals, both variables, direction, strength, and the fact that the association is linear.
  • Calling a near-zero correlation “no relationship.” That conclusion ignores curved patterns. Inspect the scatterplot and describe any curvature or other unusual features.
  • Treating association as cause. For observational data, do not write that one variable “causes,” “makes,” or “leads to” a change in the other. Describe what tends to occur together and match any causal claim to the study design.
  • Claiming a fixed result for every case. Correlation describes an overall pattern, not what must happen to each individual. Use “tend to” rather than “always” or “each extra unit causes.”
  • Skipping the verification. A wrong sign can result from arithmetic or data-entry errors. Check that the sign agrees with the direction in the scatterplot, and use LinReg on the same paired values as a numerical check.

For full-credit work, show enough calculation to make the result traceable, then interpret it in the particular context. If a graph shows a curve or an influential unusual point, say so rather than letting one numerical summary stand in for the whole data display.

Key takeaway: A sound correlation response links correctly paired data, the calculation of \(r\), a contextual description of the linear pattern, and a check that the scatterplot and study design support the claims being made.

Check Your Understanding

Use the calculation, interpretation, and critique skills from the examples. Treat each scenario as fictional.

  1. For paired observations \((1,8)\), \((2,7)\), \((3,5)\), \((4,4)\), and \((5,1)\), calculate \(r\) using centered deviations. Show the means and sums needed, then state the direction of the association.
  2. A survey finds \(r=0.68\) between weekly hours spent gardening and the number of vegetables harvested by 30 gardeners. Write a contextual interpretation and identify one claim the correlation alone cannot support.
  3. A scatterplot shows a strong inverted-U pattern and has \(r=-0.06\). Explain what \(r\) says about linear association and what it does not say about the full relationship.
  4. A student reports \(r=1.14\). Identify the error that must be present, assuming the correlation is defined, and give one practical way to check the calculation.
  5. A fictional observational study reports a negative correlation between delivery time and customer ratings. Rewrite “Longer deliveries cause customers to give lower ratings” as a careful interpretation that stays within the evidence.