Tutorials › AP Statistics › Sources of Variability in Collected Data

Investigative questions and data collection · Tutorial 156 of 1000

Sources of Variability in Collected Data

Learn why collected data vary, how three important sources of variability differ, and what a careful data collection plan can do about them.

Beginner 9 min read

What You'll Learn

  • Distinguish sampling variability from measurement error and natural variation.
  • Explain why different random samples can produce different statistics.
  • Identify random and systematic measurement error in a data collection process.
  • Recognize genuine differences among observational units and over time.
  • Describe practical ways to reduce avoidable error without claiming to remove all variability.

Why Collected Data Differ

In the previous tutorial, Choosing Between a Survey, Observational Study, and Experiment, we focused on matching a data collection method to a research question. Whatever method researchers choose, the resulting data can vary. Two reasonable samples may give different estimates, a measuring tool may not record a quantity exactly, and the people or objects being studied may genuinely differ.

These sources of variability matter because they affect what a data set can tell us. If researchers see a difference between two groups, for example, they should consider whether it reflects genuine differences, how the data were measured, or which observations happened to be selected. A numerical result by itself does not reveal the source.

Definition: Sampling variability is the variation in a statistic from one sample to another when samples are selected from the same population using the same method. Measurement error is the difference between a recorded measurement and the quantity the measurement is intended to capture. Natural variation is genuine variation among observational units or in a variable over time.

These sources can occur together. A survey statistic can vary because a different random sample was selected, because respondents do not report an experience exactly, and because people truly have different experiences. Recognizing the sources helps us describe the limits of the data without assuming every difference is a mistake.

Sampling Variability: Different Samples, Different Statistics

A sample is only part of a population. If researchers take another sample from the same population, they will usually select at least some different individuals. The resulting statistic—such as a sample mean or sample proportion—may therefore differ. This variation among sample statistics is sampling variability.

Random selection helps prevent a researcher from deliberately or accidentally favoring certain members of the population. It does not guarantee that the sample will match the population perfectly. A random sample can, by chance, include more members of one kind than another. As discussed in Generalizing Results to a Population, random selection can support generalizing to the population from which the sample was selected; it does not make a sample statistic identical to the population parameter.

A larger sample will generally show less sampling variability than a smaller sample selected in the same way. More observations provide more information about the population. But increasing the sample size does not fix a biased selection method, and it does not automatically fix inaccurate measurements. A very large sample gathered from only volunteers, for instance, can still fail to represent the population of interest.

Worked Example: Two Samples of Commute Times

Imagine a small population of eight students whose one-way commute times, in minutes, are 8, 10, 12, 14, 16, 18, 20, and 22. Two different random samples of four students are selected. Sample A has times of 8, 12, 16, and 20 minutes. Sample B has times of 10, 14, 18, and 22 minutes. Compare the sample means with the population mean.

First, the mean commute time for the full population is

$$ \mu=\frac{8+10+12+14+16+18+20+22}{8} =\frac{120}{8}=15\text{ minutes}. $$

A second check is to pair the smallest and largest values: each of the four pairs, (8, 22), (10, 20), (12, 18), and (14, 16), has mean 15. Thus the population mean is 15 minutes.

For Sample A,

$$ \bar{x}_A=\frac{8+12+16+20}{4} =\frac{56}{4}=14\text{ minutes}. $$

The values in Sample A are evenly spaced around 14: the pairs (8, 20) and (12, 16) each have mean 14. For Sample B,

$$ \bar{x}_B=\frac{10+14+18+22}{4} =\frac{64}{4}=16\text{ minutes}. $$

The pairs (10, 22) and (14, 18) each have mean 16, which checks the calculation. Both samples have four students and come from the same population, yet their sample means differ by 2 minutes. That difference is an example of sampling variability. Neither sample mean has to equal the population mean.

This example does not show that random sampling always produces means equally close to the population mean. It illustrates that different samples can produce different statistics even when the population and sample size stay the same.

Measurement Error: The Recorded Value Is Not Always Exact

Measurement error occurs when the recorded value differs from the quantity the study intends to measure. A scale might be miscalibrated, a stopwatch might be started late, a survey respondent might misremember an event, or two observers might apply a category definition differently. Errors can be small or substantial, and they can occur in numerical or categorical data.

Some measurement error is random: readings vary unpredictably, sometimes above and sometimes below the intended value. Other error is systematic: a consistent problem pushes readings in one direction. A thermometer that always reads 0.4 degrees Celsius too high has a systematic offset. Repeating the measurement may reveal inconsistent readings, but repeating it with the same miscalibrated thermometer does not remove the offset.

When a trusted reference value is available, a measurement’s error can be calculated as recorded value minus reference value. In many studies, however, the exact value is not known. Researchers may be able to reduce potential error with clear variable definitions, suitable tools, consistent procedures, training, and careful recording. These steps improve measurement; they do not promise perfect accuracy.

Worked Example: Checking a Greenhouse Thermometer

A greenhouse worker checks a thermometer against a trusted reference reading of 20.0°C. The thermometer displays 20.4°C, 20.3°C, 20.5°C, 20.4°C, and 20.4°C in five checks. Find the mean displayed value and describe the evidence about measurement error.

The mean displayed value is

$$ \bar{x}=\frac{20.4+20.3+20.5+20.4+20.4}{5} =\frac{102.0}{5}=20.4^\circ\text{C}. $$

To check the mean, compare each reading with 20.4°C: the deviations are 0, \(-0.1\), \(+0.1\), 0, and 0 degrees. They add to 0, so the mean is indeed 20.4°C. The mean reading’s error relative to the reference is

$$ 20.4^\circ\text{C}-20.0^\circ\text{C}=+0.4^\circ\text{C}. $$

The thermometer’s average reading is 0.4°C higher than the reference. The readings also vary slightly from check to check: their range is \(20.5-20.3=0.2^\circ\text{C}\). The repeated readings show some measurement variability as well as a consistent positive offset in this set of checks. The worker should investigate calibration before relying on the thermometer. Simply taking more readings with the same instrument might give a more stable average display, but would not by itself correct a systematic offset.

Natural Variation: Real Differences in the World

Not every difference among data values is an error. Observational units may genuinely differ. For example, plants grown in the same garden can reach different heights because of biological differences and small differences in their growing conditions. The heights vary because the plants vary; a ruler can measure those real differences accurately.

Natural variation can also occur within one unit over time. A person’s pulse may change from one moment to another, even when the measurement procedure is consistent. A question about a person’s pulse therefore needs to specify when or under what conditions it is measured. Otherwise, readings taken at different times may reflect genuine changes in the person’s pulse as well as measurement error.

In practice, “natural variation” does not always mean that researchers know the exact cause of each difference. It means the values may reflect real differences in the units or real changes over time, rather than a flaw in the selection or measurement process. As with measurement error, careful definitions and consistent procedures help researchers interpret what the data represent.

Worked Example: Heights of Garden Seedlings

A student measures five seedlings grown in the same garden bed. Their heights are 12.1, 12.8, 13.4, 11.9, and 14.0 centimeters. Find their mean height and range, then explain whether the spread must indicate measurement error.

The mean height is

$$ \bar{x}=\frac{12.1+12.8+13.4+11.9+14.0}{5} =\frac{64.2}{5}=12.84\text{ cm}. $$

As a check, compare each value with 12.8 cm. The deviations are \(-0.7\), 0, \(+0.6\), \(-0.9\), and \(+1.2\) cm. Their sum is \(0.2\) cm, so the mean is \(12.8+(0.2/5)=12.84\) cm. The range is

$$ 14.0-11.9=2.1\text{ cm}. $$

The tallest and shortest seedlings differ by 2.1 cm. That spread does not, by itself, show that the student measured incorrectly. It could reflect natural differences in seedling growth. A ruler that was used inconsistently could add measurement error, too; the summary statistics alone cannot separate the sources. A careful conclusion is that the observed heights vary, and the variation may include genuine differences among seedlings as well as any measurement error in the procedure.

Separating Sources and Improving a Data Collection Plan

A useful way to think about a data collection plan is to ask three questions. First, who or what will be selected? That points to sampling variability and possible selection problems. Second, how will each variable be measured? That points to measurement error. Third, how much might the observational units genuinely differ, or the variable change over time? That points to natural variation.

These are related but not interchangeable. A biased sample is not simply another name for sampling variability. Sampling variability describes how statistics differ across samples selected by the same method; bias is a systematic problem that makes the method tend to miss or misrepresent part of the population. Similarly, a wide spread in the data does not prove poor measurement. It may reflect real differences among units.

Consider an observational study of the time residents spend waiting for a bus. A random sample of residents may produce different average waiting times on different occasions. That is sampling variability. Residents may start their timers late or remember a past wait inaccurately; those are possible measurement errors. Buses may genuinely arrive at different intervals and residents may arrive at different points in the schedule; those differences contribute natural variation. A data collection plan can address each source differently: choose an appropriate sample, define and measure waiting time consistently, and record relevant context such as time of day.

The method should fit the question, as explained in the previous tutorial. A survey may be appropriate for a reportable experience, while direct timing may be better for an observed wait. Neither choice guarantees that all error disappears. Researchers should explain what they measured and recognize limits that remain.

Common Mistakes and AP Exam Tips

  • Calling every difference “measurement error.” Real differences among people, objects, or times can produce different values. For full credit, say whether the difference could reflect natural variation, measurement error, or both, and connect the claim to the context.
  • Assuming random selection produces the same statistic every time. Random selection avoids choosing only preferred cases, but different random samples can still produce different statistics. State that this difference is sampling variability.
  • Assuming a larger sample fixes every problem. A larger sample generally reduces sampling variability, but it cannot automatically repair a biased sample or a systematic measuring-tool problem. Identify which source the proposed change addresses.
  • Confusing a consistent offset with random fluctuation. Readings that repeatedly fall above a trusted reference suggest a possible systematic measurement error. Readings that vary around a reference may indicate random measurement variation, though the process should still be checked.
  • Treating observed spread as proof of a flawed study. Spread can be an expected feature of real data. Describe what varies and avoid claiming a cause unless the data collection process provides evidence for it.

A strong AP response names the source and explains its role in context. For example: “Different random samples of residents could produce different sample mean waiting times, so the estimate has sampling variability. Inaccurate timing could add measurement error, while actual differences in bus arrival intervals could create natural variation.” That wording distinguishes the sources instead of treating all variability as a single problem.

Key takeaway: Collected data can vary because different samples produce different statistics, measurements do not perfectly capture the intended quantity, or observational units and conditions genuinely differ. Identify the source that fits the situation; improving one part of a data collection plan does not automatically remove the others.

Check Your Understanding

For each situation, identify the most relevant source or sources of variability and briefly explain your reasoning.

  1. A random sample of 40 residents gives a different estimate of weekly library visits from another random sample of 40 residents.
  2. A bathroom scale displays 0.5 kilograms too much each time it is checked against a known reference weight.
  3. Two trees of the same species, grown in the same park, have different measured heights.
  4. A student answers a survey question about how many hours they slept last week but cannot remember the exact number.
  5. Explain why increasing a sample size might help with one source of variability but not necessarily solve all problems in a study.