Tutorials › AP Statistics › Undercoverage Bias

Random sampling · Tutorial 172 of 1000

Undercoverage Bias

Identify who is missing from a sampling frame, explain how those omissions may affect results, and state which population a sample can reasonably describe.

Beginner 8 min read

What You'll Learn

  • Define undercoverage bias and identify which population members a sampling frame leaves out.
  • Distinguish omissions from the sampling frame from people selected but not responding.
  • Explain why a random sample from an incomplete frame may still misrepresent the population.
  • Use a constructed example to show how omitted groups can change a population proportion.
  • Describe when undercoverage may or may not bias a particular survey result.
  • Limit conclusions to the group the sampling frame actually covers.

When the Frame Leaves People Out

In Population, Sampling Frame, and Sample, you learned to distinguish the population of interest from the sampling frame used to select a sample. A sampling frame is the list or source from which individuals are selected. If that frame leaves out part of the population, even a carefully selected sample may not represent everyone the study is about.

For example, a survey about all adults in a town might use a list that contains only people with landline telephone numbers. Adults who have only mobile phones are part of the population of interest, but they are missing from the frame. Randomly selecting names or numbers from the available list does not give those omitted adults a chance to be selected.

Definition: Undercoverage occurs when some members or groups in the population of interest are missing from the sampling frame or are represented less adequately than others. Undercoverage bias can result when the omitted or underrepresented members differ from the rest of the population on the variable being measured, causing the sample’s results to systematically misrepresent the population.

The key diagnostic is to compare the population you want to describe with the people or units the frame can actually reach. Ask: Who is included? Who is missing? Could the missing group answer the survey differently from those included? Those questions connect the frame’s coverage to the study’s specific variable.

Random Selection Cannot Include Someone Missing From the Frame

As discussed in Why Random Selection Matters, chance-based selection reduces the opportunity for a researcher’s preferences to determine who enters a sample. But chance can select only from the frame it is given. If a group is absent from that frame, its members have no chance of selection. A simple random sample of the frame can therefore be genuinely random and still fail to represent the full population of interest.

This is a different concern from ordinary sampling variability, introduced in Sources of Variability in Collected Data. Sampling variability means that different random samples from the same frame can produce different statistics. Undercoverage is a problem with the frame itself: one or more groups are missing or poorly represented before the sample is selected. Increasing the sample size can reduce the effect of random fluctuation, but it does not add omitted groups to the frame.

Undercoverage also differs from the voluntary response bias discussed in the previous tutorial. With undercoverage, people are absent from the source used to select the sample. With voluntary response, an open invitation allows people to choose themselves into the sample. In a survey, more than one concern can occur, but identifying the frame’s omissions is a separate step.

Diagnostic questions: Name the population of interest. Then describe the sampling frame and identify who cannot appear in it or is less likely to be represented. Finally, consider whether those groups might differ on the variable being measured.

Worked Example: A Town Survey About Bus Service

Worked Example: Who Is Missing From a Landline List?

Imagine a fictional town with 1,000 adult residents. The town wants to estimate the proportion who support extending bus service. Its survey team selects a random sample from a list of residents with landline telephone numbers. The list covers 600 adults. In a constructed illustration, 360 of those 600 adults support the extension. The 400 adults not on the list include many residents who have only mobile phones; 120 of them support the extension.

State: The population of interest is all 1,000 adults in the town. The variable is whether an adult supports extending bus service. The sampling frame covers only the 600 adults with landline numbers and omits 400 adults.

Plan: A random sample from the landline list can represent that list if the selection is carried out properly, but it cannot directly represent adults missing from the list. Undercoverage could bias the townwide estimate if support for bus service differs between adults with landlines and adults without them.

Do: The support proportion among adults covered by the frame is

$$ \frac{360}{600}=0.60 $$

The support proportion among the omitted adults is

$$ \frac{120}{400}=0.30 $$

In the constructed population, 360 covered adults and 120 omitted adults support the extension, for 480 supporters out of 1,000 adults:

$$ \frac{360+120}{1000}=\frac{480}{1000}=0.48 $$

The frame’s support proportion is 0.60, which is 0.12, or 12 percentage points, higher than the constructed townwide proportion of 0.48.

Conclude: In this illustration, the landline frame undercovers adults without landlines, and that omitted group has lower support for extending bus service. As a result, a sample that reflects the frame could overestimate support among all town adults. The 0.60 proportion describes the covered group; by itself, it does not establish the proportion among all 1,000 adults.

Check Whether the Missing Group Matters for the Variable

Finding an omission is an important warning, but it does not automatically tell you the direction or size of bias. The result depends on how the omitted group differs on the variable of interest. If omitted people tend to have higher values, a frame-based estimate might be too low. If they tend to have lower values, it might be too high. If their distribution is similar to that of the covered group, the omission may have little effect on that particular estimate.

For example, suppose a survey’s frame omits some residents who lack internet access. That omission is clearly a coverage problem for a survey of all residents. If the variable is preferred method for receiving online notices, omitted residents may differ substantially from those on the frame. If the variable is a characteristic unrelated to internet access, the omission may have less effect on its estimate. The frame’s quality should therefore be evaluated in relation to the study question, not in the abstract.

In practice, the population value for omitted people is usually unknown—that is why the survey is being conducted. You should not claim that undercoverage definitely raised or lowered an estimate unless the scenario provides evidence about how the omitted group differs. It is often more accurate to say that undercoverage could bias the estimate and explain why the group might differ.

Worked Example: A School Technology Survey

A fictional district wants to estimate the proportion of all its high-school students who would use a new online tutoring service. It selects a random sample from the district’s student-app roster. The roster includes 700 of the district’s 1,000 high-school students. In a constructed illustration, 140 of the 700 students on the roster say they would use the service. The 300 students missing from the roster include many students whose families have not activated the app; 180 of those omitted students say they would use the service.

State: The population of interest is all 1,000 high-school students in the district. The variable is whether a student would use the online tutoring service. The frame is the student-app roster, which omits 300 students.

Plan: Randomly sampling from the roster does not give students missing from it a chance to be selected. If app access or family circumstances are related to interest in online tutoring, the frame may underrepresent students with different levels of interest.

Do: Among students included in the frame, the proportion who would use the service is

$$ \frac{140}{700}=0.20 $$

Among the omitted students, the proportion in the constructed illustration is

$$ \frac{180}{300}=0.60 $$

The townwide proportion in this constructed example is

$$ \frac{140+180}{1000}=\frac{320}{1000}=0.32 $$

The frame proportion of 0.20 is 0.12, or 12 percentage points, lower than the constructed population proportion of 0.32.

Conclude: In this example, the roster omits students who have a higher rate of interest in the service. A sample drawn from the roster could therefore underestimate interest among all district high-school students. The survey can describe students represented by the roster, but generalizing to all students requires addressing the omitted group.

Frame Boundaries Can Be Easy to Miss

A frame does not have to be a printed list. It can be a database, membership roster, collection of records, or source used to identify potential participants. It may also be tied to a particular location or time. For example, asking people who enter a recreation center on weekday mornings about all city residents’ exercise habits excludes residents who do not visit that center and those who visit at other times. The source of recruitment defines who can enter the sample.

A frame may also be out of date. A customer list that omits new customers, a school roster that has not been updated, or a directory that retains people who have moved can fail to match the population of interest. The issue is not simply whether the frame is called a list; it is whether the frame adequately covers the defined population for the time and place of the study.

This is why Defining the Population of Interest comes before evaluating a frame. A frame can be suitable for one question and inadequate for another. A list of current library-card holders might cover the population of card holders well, but it would not cover all residents if the study question is about every resident’s views on library funding.

Worked Example: A Community Center Feedback Survey

A fictional city wants to estimate how satisfied all residents are with its public parks. Staff place feedback cards at a community center and ask visitors to complete them. The center is used by some residents, but the city has no list of all residents who visit parks, and residents who do not use the center cannot receive a card there.

State: The population of interest is all city residents. The variable is each resident’s satisfaction with the public parks. The feedback-card source reaches community-center visitors who notice and choose to complete a card; it does not cover all city residents.

Plan: The source may underrepresent residents who do not use the center, including people who use different parks or do not visit the center at all. If center visitors have different experiences or satisfaction levels, the responses may not reflect all residents. Since completing a card is optional, self-selection may also affect who responds.

Do: Suppose 72 of the 90 returned cards say that the respondent is satisfied. The satisfaction proportion among card respondents is

$$ \frac{72}{90}=0.80 $$

Thus, 80% of the people who returned cards reported satisfaction. The city has no corresponding information from this feedback method about residents outside its reach or about people who visited but did not return a card. The 80% calculation is valid for the returned cards, but it does not calculate the proportion of all residents who are satisfied.

Conclude: The city should report that 80% of the returned feedback cards indicated satisfaction, not that 80% of all residents are satisfied. The feedback source undercovers residents who do not visit the center, and optional participation creates an additional concern. The result can summarize the respondents, but it does not support a reliable generalization to all city residents.

Common Mistakes and AP Exam Tips

  • Assuming random selection fixes an incomplete frame. Random selection only operates on the frame. A full-credit response names the group missing from the frame and explains that those individuals had no chance to be selected.
  • Calling every difference from the population “undercoverage.” Identify the actual mechanism. Undercoverage concerns who is missing or poorly represented in the frame; voluntary response concerns people choosing themselves into a sample.
  • Claiming the direction of bias without evidence. The missing group could have higher or lower values on the measured variable. Explain a plausible connection when the context supports one, or say the direction cannot be determined from the information given.
  • Assuming any omission must strongly change every result. Omitted groups are most concerning when they may differ on the variable being measured. Explain that connection rather than treating all frame defects as equally important for every question.
  • Generalizing a frame-based statistic to everyone in the population. State whom the frame covers and whom the result directly describes. A broader population claim needs a frame and selection method that adequately cover that population.
  • Thinking that a larger sample repairs undercoverage. A larger sample can provide more information about the frame, but it cannot include a group absent from the frame. The coverage problem must be addressed in the sampling plan.

For a strong AP response, identify the population of interest, describe the sampling frame, name the omitted group, and connect the omission to the measured variable. Then explain how this limits generalization. Use cautious wording such as “may overestimate” or “could be biased” unless the scenario provides enough information to establish a particular direction.

Key takeaway: Undercoverage occurs when a sampling frame omits or poorly represents part of the population of interest. Random selection from that frame cannot include people who are missing. Undercoverage can bias results when omitted groups differ on the variable being measured, so describe who the sample can represent and limit conclusions accordingly.

Check Your Understanding

For each situation, identify the population, the frame’s coverage problem, and what the results can support.

  1. A survey about all residents’ use of a city website selects a random sample from the list of residents who have created website accounts. Identify one group that may be undercovered and explain why random selection from the account list does not reach them.
  2. A college surveys current students using its enrollment list but wants to draw a conclusion about everyone who applied to the college. Explain the mismatch between the population of interest and the frame.
  3. A survey frame omits a group, but the scenario gives no information about how that group differs on the measured variable. What can you say about the possibility and direction of bias?
  4. A random sample from a frame gives an estimate of 0.42 for a population proportion. Explain why a larger sample from the same incomplete frame would not necessarily solve undercoverage.
  5. A library asks people visiting one branch on weekday afternoons to rate all city library services. Identify a possible coverage limitation and write a cautious sentence describing whom the responses represent.