Tutorials › AP Statistics › Constructing a Scatterplot by Hand

Scatterplots and association · Tutorial 802 of 1000

Constructing a Scatterplot by Hand

Practice turning ten paired measurements into a clear, correctly scaled scatterplot by hand.

Intermediate 9 min read

What You'll Learn

  • Choose axis ranges that include all observations and make the pattern easy to see.
  • Select convenient, evenly spaced intervals for each axis.
  • Label axes with the variable names and units.
  • Plot each observation as one correctly ordered \((x,y)\) point.
  • Check a completed scatterplot for missing, misplaced, or distorted points.

From Paired Measurements to Plotted Points

In “Explanatory and Response Variables in Scatterplots,” you learned to use the research question to assign the explanatory variable to the horizontal \(x\)-axis and the response variable to the vertical \(y\)-axis. Now you can use that decision to construct a scatterplot by hand. The main tasks are choosing useful axis ranges, marking a consistent scale, and plotting each pair without separating its measurements.

A scatterplot displays each observational unit as one point. If a plant receives \(x\) hours of light and grows to \(y\) centimeters tall, its point is \((x,y)\): move to \(x\) on the horizontal axis, then to \(y\) on the vertical axis. Plotting the values in the opposite order would show a different point and would not represent the recorded pair correctly.

Definition: A scatterplot is a display of paired values for two quantitative variables. Each observational unit contributes one point \((x,y)\), with the explanatory variable on the horizontal axis and the response variable on the vertical axis.

Choosing Axis Ranges and Scales

Start by finding the smallest and largest observed values for each variable. Choose a range on each axis that includes every observation. Then select a convenient interval between tick marks, such as 1, 2, 5, 10, or a suitable decimal amount. The intervals must be equal along a given axis: if the marks represent 0, 2, 4, and 6, each equal distance on the page represents 2 units.

The two axes do not have to use the same interval or begin at zero. Their variables may have different units and different ranges. For example, one axis might increase by 5 degrees at each tick while the other increases by 10 kilowatt-hours. Starting near the data can make the points easier to distinguish, provided the range includes all observations and the scale is clearly labeled.

Guidelines for useful scales: Include the full range of observed values; use equal spacing for equal numerical increments on each axis; choose intervals that are easy to count; and label each axis with the variable name and units. A scale should use enough of the plotting area to make the points readable without crowding them.

A range that is far wider than the data can compress the points into a small part of the graph. A range that is too narrow can leave out observations or make the display difficult to read. Do not change the spacing partway along an axis to fit the data. If you choose a scale that does not start at zero, make the starting value obvious; do not draw a break that might make readers misread the scale.

Before drawing points, make the axes and tick marks. A straightedge can help keep the axes neat, but the important feature is a consistent, readable scale. Add a clear label and units to each axis; a brief title can also identify what the display shows. Then plot the pairs one at a time. A light pencil mark can help you check a point’s horizontal and vertical locations before making the dot.

1
Confirm the variables.
Use the research question to identify \(x\), the explanatory variable, and \(y\), the response variable.
2
Check the observed ranges.
Find the minimum and maximum for each variable so every point will fit on the graph.
3
Draw and label the axes.
Choose convenient, evenly spaced intervals and write the variable names and units beside the correct axes.
4
Plot and audit the pairs.
For each row, locate \(x\) first and \(y\) second. Check that every observational unit has exactly one point in the correct location.

A scatterplot uses separate points, not a line connecting one observation to the next. Connecting points can suggest an ordered path or values between observations that the data do not provide. If two observational units have the same values, their points overlap; in that case, note the overlap rather than shifting a point to a different coordinate.

Worked Example: Plotting Plant Growth and Light Exposure

A student records daily light exposure and plant height for ten fictional seedlings of the same variety. The question is whether light exposure helps explain differences in height. Light exposure is \(x\), measured in hours, and height is \(y\), measured in centimeters.

SeedlingLight exposure (hours)Height (cm)
A16
B28
C310
D410
E512
F614
G714
H818
I918
J1020

Choose the ranges. The light values run from 1 to 10 hours, and the heights run from 6 to 20 cm. Use an \(x\)-axis from 0 to 10 hours marked every 1 hour and a \(y\)-axis from 0 to 20 cm marked every 2 cm. These ranges include all observations and make the values easy to locate.

Label and plot. Label the horizontal axis “Light exposure (hours)” and the vertical axis “Height (cm).” For Seedling A, move to \(x=1\) and then \(y=6\) to plot \((1,6)\). For Seedling D, plot \((4,10)\). Continue in the same way for each row, keeping the two measurements for each seedling together.

Check the plotted locations. The \(x\)-coordinates are the integers 1 through 10, and the \(y\)-coordinates are 6, 8, 10, 10, 12, 14, 14, 18, 18, and 20. Some seedlings share a height, so their points align horizontally, but they do not overlap because their light values differ.

The following grid shows the intended positions. The horizontal labels are hours, and the row labels are height in centimeters. A dot represents one seedling.

Height \ Hours012345678910
20●
18●●
16
14●●
12●
10●●
8●
6●
4
2
0

Counting the dots gives ten plotted observations, matching the ten seedlings. The display suggests that greater light exposure tends to occur with greater height in this invented set of data. The scatterplot shows an association in these observations; by itself, it does not establish that light exposure caused the differences.

Worked Example: Selecting Ranges When Neither Variable Starts Near Zero

A fictional community energy team records afternoon humidity and electricity demand for ten days. The question is whether humidity is associated with demand. Humidity is the explanatory variable, \(x\), measured as a percentage; demand is the response variable, \(y\), measured in kilowatt-hours.

DayHumidity (%)Electricity demand (kWh)
142108
245115
348111
452124
555127
658132
761129
865143
969148
1073154

Find the ranges and choose intervals. Humidity ranges from 42% to 73%, so use an \(x\)-axis from 40% to 75%, marked every 5 percentage points. Demand ranges from 108 to 154 kWh, so use a \(y\)-axis from 100 to 160 kWh, marked every 10 kWh. Neither axis needs to start at zero: both chosen ranges include all the observations and give the points room on the graph.

Plot values between tick marks accurately. Label the axes “Afternoon humidity (%)” and “Electricity demand (kWh).” For Day 1, humidity 42% is two-fifths of the way from the 40% tick to the 45% tick, and 108 kWh is just below the 110 kWh tick. Plot the point at those two positions. For Day 8, plot \((65,143)\): the \(x\)-value is exactly at the 65% tick, while the \(y\)-value is three-tenths of the way from 140 to 150 kWh.

Finish and check. Plot all ten rows in this manner. Values such as 48% and 111 kWh fall between labeled marks, so estimate their positions using the consistent scale rather than rounding them to a nearby tick. Confirm that the plotted point for each day uses that day’s humidity and demand together.

This example shows why the two axes can use different intervals and different starting values. What matters is that each scale is uniform, every observation is included, and the labels make the units clear. A range that begins at the observed minimum without a little margin can make the end points look cramped; here, 40% and 100 kWh provide a readable margin below the data.

Worked Example: Auditing a Scatterplot of Wait Time and Rating

A fictional service center records the wait time and customer rating for ten visits. Wait time is the explanatory variable, \(x\), in minutes. Rating is the response variable, \(y\), on a scale from 1 to 5. The goal is to make a hand plot and check that it represents every visit correctly.

VisitWait time (minutes)Rating (points)
A84.8
B124.5
C174.4
D214.1
E274.0
F313.8
G363.6
H433.5
I493.2
J563.0

Set up the graph. Use a horizontal axis from 0 to 60 minutes, marked every 10 minutes, and a vertical axis from 2.5 to 5.0 rating points, marked every 0.5 point. Label the axes “Wait time (minutes)” and “Customer rating (points).” The vertical range includes the lowest rating, 3.0, and the highest, 4.8, while making differences among these ratings visible.

Plot several coordinates carefully. Visit A contributes \((8,4.8)\): place it between 0 and 10 minutes, near 10, and near the top of the rating scale. Visit E contributes \((27,4.0)\): place it between 20 and 30 minutes, closer to 30, at the 4.0 rating mark. Visit J contributes \((56,3.0)\): place it between 50 and 60 minutes, closer to 60, at the 3.0 mark. Use the same method for the other seven visits.

Audit the result. There should be ten points, with no line joining them. Check each coordinate against its row: for example, the point at 56 minutes must have rating 3.0, not 5.6 or 56. The plotted values tend to show lower ratings at longer wait times, but the hand plot alone does not justify a causal claim about wait time.

The rating axis does not begin at zero because the data range from 3.0 to 4.8 and the purpose is to display their differences clearly. Its lower endpoint is labeled, the tick spacing is constant, and all values fit. A reader can therefore interpret the scale rather than mistaking the display for one that covers a different range.

Common Mistakes and AP Exam Tips

  • Reversing the coordinates: A pair \((x,y)\) means horizontal value first, vertical value second. Check which variable is explanatory before plotting.
  • Using uneven tick spacing: Equal physical distances along an axis must represent equal numerical changes. Do not make the interval between 10 and 20 look different from the interval between 20 and 30.
  • Leaving out a point at the edge: Check the minimum and maximum in both variables before choosing endpoints. Every observed pair must fit within the graph.
  • Rounding a value to a tick mark: If a measurement lies between labeled ticks, estimate its position between them. Do not move it to a different coordinate just to make plotting easier.
  • Connecting the dots: A scatterplot represents paired observations as separate points. Do not join them in data-table order unless a different kind of display is specifically called for.
  • Forcing both axes to start at zero: Zero is not required if it is not useful for the data. Choose a clear range that includes all observations, and label a nonzero starting point unmistakably.
  • Forgetting units or changing the pairing: Label each axis with its variable and units, and plot each row as one pair. Mixing values from different observational units changes the data.

For full-credit communication, state which variable is on each axis, give the ranges and intervals used, and describe how each pair becomes a point. If explaining a completed plot, check that its scales are even, its labels include units, and its plotted points match the data.

Key takeaway: A clear hand-drawn scatterplot includes all paired observations, places the explanatory variable on the horizontal axis and the response variable on the vertical axis, and uses labeled, evenly spaced scales chosen to show the data clearly.

Check Your Understanding

Answer each question using the plotting principles from this tutorial.

  1. A data set has \(x\)-values from 14 to 37 and \(y\)-values from 102 to 169. Suggest a convenient range and tick interval for each axis that includes all the data.
  2. For one observational unit, \(x=24\) and \(y=7.5\). Describe the order of movements you would make to plot its point.
  3. Why can the horizontal and vertical axes use different tick intervals? What must be consistent within each axis?
  4. A student makes a graph with \(x\)-ticks at 0, 5, 10, and 20, each separated by the same physical distance. What is wrong with that scale?
  5. Why might a scatterplot use an axis range that starts above zero, and what should the graph show clearly when it does?