Tutorials › AP Statistics › Writing a Correlation Conclusion on the AP Exam

Correlation · Tutorial 839 of 1000

Writing a Correlation Conclusion on the AP Exam

Practice writing precise conclusions about \(r\) that describe the observed linear association in context and respect what the study design can establish.

Intermediate 9 min read

What You'll Learn

  • Use a repeatable structure to interpret the direction and strength of a reported correlation.
  • Name the individuals, both quantitative variables, and relevant units in a conclusion.
  • Explain why an interpretation of \(r\) should identify the association as linear.
  • Connect the strength description to the scatterplot instead of relying on a universal cutoff.
  • Add a causation caveat that matches the study design and the claim being evaluated.
  • Revise vague, incomplete, or overstated interpretations into AP-style conclusions.

From a Number to a Complete Conclusion

In “Common Misinterpretations of \(r\),” you learned to keep correlation distinct from a percent, a slope, and a causal effect. Now the goal is to use \(r\) accurately in a complete AP-style conclusion. A number by itself is not a conclusion: your writing should tell a reader what the correlation says about the variables in the particular situation, and what it does not establish.

As in “Interpreting a Correlation in Context,” name both quantitative variables and describe their direction and strength. For a complete response, also make clear that \(r\) summarizes a linear association, identify the individuals represented by the data, and avoid suggesting that one variable causes the other to change unless the study design supports a causal conclusion.

Conclusion checklist: Name the individuals, the two quantitative variables and their units; describe the direction and strength of the linear association in context; and state a causation caveat that fits the study design.

A useful sentence pattern is: “Among [individuals], there is a [strength] [positive or negative] linear association between [explanatory variable] and [response variable]: individuals with [higher or lower explanatory-variable values] tend to have [higher or lower response-variable values]. Because [relevant design detail], this association alone does not show that [causal claim].”

This is a guide, not a script to copy mechanically. Include details that matter in the question, and do not add claims the data cannot support. The word “linear” matters because, as covered in “Correlation Measures Only Linear Association,” \(r\) describes how closely the data follow a straight-line pattern. It does not summarize every possible relationship between the variables.

Build the Interpretation One Part at a Time

Begin by identifying who or what each observation represents. These are the individuals in the data: for example, students, plants, or days. Naming them tells the reader whose measurements the conclusion describes. If a question gives only a sample, keep the conclusion about that sample unless an appropriate inference procedure supports a broader claim.

Next, name both variables, using context and units where they help identify what was measured. Avoid writing only “\(x\) and \(y\)” or “the two variables.” A reader should be able to tell whether the conclusion concerns, for example, weekly practice hours and a performance score, rather than some other measurements on the same students.

Then describe the direction and strength of the linear association. Use the sign of \(r\) for direction. Use a strength word only as a reasonable description of the particular data: a scatterplot can help you judge how closely the points follow a straight-line pattern. There is no single cutoff that makes every correlation “weak,” “moderate,” or “strong” in every setting. In “Judging Strength of an Association,” you learned to assess strength from the scatterplot; that visual check is especially useful before choosing a word.

Finally, make sure your sentence describes a tendency across the observations, not a guaranteed result for every individual. “Students with more practice hours tend to have higher scores” describes a pattern. “Every student who practices more earns a higher score” makes a claim about every case that the correlation does not establish.

1
Set the scope.
Name the individuals represented and, when the data are a sample, keep the conclusion about those observations.
2
Name both variables.
Use their contextual names and units as appropriate. Keep the explanatory and response variables in the roles established by the research question.
3
Describe the linear pattern.
Use the sign of \(r\) for direction and the scatterplot to help describe strength. State that the association is linear.
4
Check the causal wording.
Say what the study design allows. An observed correlation alone does not establish that changing one variable causes a change in the other.

Worked Examples: Writing and Revising Conclusions

Worked Example: Practice Time and a Performance Score

A fictional observational data set records weekly practice time, in hours, and a performance score, in points, for 32 members of a youth music program. The reported correlation is \(r=0.76\). The scatterplot shows a fairly tight, roughly straight upward pattern with no point far from the overall pattern. Write a complete interpretation.

The individuals are the 32 program members. The variables are weekly practice time and performance score. Since \(r\) is positive, the association is positive: higher practice times tend to go with higher scores. The scatterplot supports describing the linear association as fairly strong, because the points follow the upward straight-line pattern fairly closely.

A suitable conclusion is: “Among these 32 youth music program members, there is a fairly strong positive linear association between weekly practice time, in hours, and performance score, in points. Members who practiced more hours per week tended to have higher performance scores. Because these are observational data, the association alone does not show that practicing more caused members to earn higher scores.”

The final sentence is a caveat that matches the design: practice time was recorded, not assigned by researchers. As discussed in “Lurking Variables and Confounding,” other factors could be related to both practice time and performance. The conclusion does not claim that any particular factor explains the pattern; it simply avoids treating the correlation as proof of cause and effect.

Worked Example: Repair an Incomplete Interpretation

A fictional school survey records students’ commute distance, in kilometers, and the number of days they arrived late during a month. For 45 surveyed students, \(r=0.41\). The scatterplot is roughly linear and shows a positive trend with noticeable scatter. A student writes, “The correlation is 0.41, so longer commutes make students late.” Revise this response into a complete, careful interpretation.

The draft gives the numerical value but does not explain its direction and strength in context. It also uses “make,” which claims causation. The individuals are the surveyed students; the variables are commute distance and late-arrival days. The positive value and upward pattern indicate that students with longer commutes tended to have more late-arrival days. Given the noticeable scatter, “moderate” is a reasonable description of this particular linear association.

A revised conclusion is: “Among the 45 surveyed students, there is a moderate positive linear association between commute distance, in kilometers, and the number of days late during the month. Students with longer commutes tended to have more late-arrival days. Because the survey recorded existing commutes and lateness rather than randomly assigning commute distances, these data alone do not show that a longer commute caused a student to arrive late more often.”

Notice what the revision does not say. It does not claim that every student with a longer commute was late more often, that lateness rose by 0.41 days per kilometer, or that the result necessarily applies to all students at the school. Those claims are not interpretations of the reported correlation.

Worked Example: A Small Correlation with a Curved Pattern

A fictional environmental data set records afternoon temperature, in degrees Celsius, and the number of visitors to a public greenhouse on 28 days. The reported correlation is \(r=0.09\). The scatterplot, however, has a clear curved pattern: visitor counts tend to be higher on mild days and lower on the coolest and hottest days. Write a conclusion that interprets \(r\) without overlooking the graph.

The individuals are the 28 days. The correlation is close to zero, so it indicates little linear association between temperature and visitor count in these observations. But the scatterplot shows a noticeable curved relationship. Therefore, it would be wrong to conclude that temperature and visitor count have no relationship at all. The correlation summarizes only their linear association.

A careful response is: “For these 28 days, there is little linear association between afternoon temperature, in degrees Celsius, and the number of greenhouse visitors. However, the scatterplot shows a curved pattern, with more visitors on mild days than on the coolest or hottest days, so the small value of \(r\) does not mean that there is no relationship. These observed data alone do not establish that temperature caused the changes in visitor counts.”

This example shows why checking form is part of writing well about \(r\). The value \(0.09\) is not incorrect; it answers a limited question about linear association. The graph reveals a feature that a correlation conclusion must not erase.

Match the Causation Caveat to the Study

A causation caveat should be accurate, not automatic filler. For observational data, a careful conclusion commonly says that the association alone does not show that changing the explanatory variable caused a change in the response variable. This is the distinction emphasized in “Why Correlation Does Not Imply Causation.” Naming a possible lurking variable can help explain why the causal claim is uncertain, but do not present a possible explanation as a proven one.

Study design matters. If researchers randomly assign a treatment in an experiment, the design may provide evidence about cause and effect. Even then, do not claim that the correlation coefficient itself proves causation. Describe what was measured and assigned, and make causal claims only to the extent justified by the experiment. In many correlation questions, the data are observational, so a brief, explicit caveat is appropriate.

Key distinction: Describe the observed association with \(r\). Evaluate a causal claim using the study design. For observational data, the correlation alone does not establish that one variable caused the other to change.

Common Mistakes and AP Exam Tips

  • Writing only the number. “\(r=0.76\)” does not interpret the result. State the direction and strength of the linear association and name the variables in context.
  • Leaving out “linear.” A correlation is not a summary of every possible pattern. Use “linear association,” and check the scatterplot for curvature or other unusual features.
  • Leaving out the individuals or units. “There is a positive association” is vague. Say whose observations are involved and identify what each variable measures. Include units when they clarify the measurements.
  • Using causal language for observed data. “Practice time increases scores” can sound like a cause-and-effect claim. For observational data, describe a tendency—such as “more practice time tended to go with higher scores”—and state that the association alone does not establish causation.
  • Making a claim about every case. A correlation describes an overall pattern, not a guaranteed outcome for every individual. Use “tend to” rather than “always.”
  • Overstating the population reach. If the question supplies results for a sample, do not silently turn the conclusion into a claim about everyone in a broader population.
  • Letting the strength label replace evidence. “Strong” or “moderate” should fit the scatterplot. Do not rely on a universal numerical cutoff, and do not confuse strength with steepness.

For full-credit communication, make each part of the conclusion do a clear job: context identifies the data, direction and strength describe the observed linear pattern, and the caveat keeps the claim within the limits of the design. A concise answer can still be complete if it includes those elements and does not add unsupported claims.

Key takeaway: A complete AP-style interpretation of \(r\) names the individuals and variables, describes the direction and strength of the linear association in context, and avoids a causal claim that the study design does not support.

Check Your Understanding

For each situation, identify what a complete interpretation should include and what claim it should avoid.

  1. A fictional survey of 36 gardeners finds \(r=-0.63\) between weekly watering time, in minutes, and the number of wilted leaves on a plant. Write a context-based interpretation that includes direction, strength, linearity, and a suitable caveat.
  2. A student interprets \(r=0.52\) for two measured variables by saying, “There is a 52% increase in the response.” Explain what is missing or incorrect, then describe what the value of \(r\) can support.
  3. A scatterplot shows a pronounced U-shaped pattern and \(r=0.04\). What can you say about the linear association, and why would “there is no relationship” be an overstatement?
  4. A fictional observational study finds a negative association between daily delivery time and customer ratings. Write a sentence that describes the pattern without claiming that longer delivery times caused lower ratings.
  5. Why should a conclusion based on a sample of surveyed students avoid automatically claiming that the same association holds for all students in the region?