Tutorials › AP Statistics › Mixed Practice: Sampling Methods and Bias

Random sampling · Tutorial 180 of 1000

Mixed Practice: Sampling Methods and Bias

Practice recognizing sampling methods and bias, then use those ideas to build and critique sampling plans.

Beginner 10 min read

What You'll Learn

  • Distinguish simple random, stratified, cluster, and systematic random sampling from convenience and voluntary response samples.
  • Support a method identification with the feature that determines how individuals were selected.
  • Trace a sampling flaw to undercoverage, nonresponse, voluntary response, or convenience selection.
  • Explain how a flaw could affect which members of the population are represented.
  • Write a reproducible plan that names a population, frame, sample size, and chance-based selection steps.
  • State what group a sample can represent, given its sampling frame and selection method.

A Mixed-Practice Routine

In Describing a Random Sampling Plan in Context, you practiced making a chance-based plan detailed enough for another person to carry out. This tutorial brings together the sampling methods and sources of bias from earlier in this unit. The goal is not just to name a method or a flaw: it is to point to the detail in the scenario that supports your answer, explain its consequence, and, when asked, propose a plan that fits the study question.

For method-identification questions, focus on what the chance mechanism selects. Does it select individuals from the full frame, some individuals from every subgroup, whole groups, or every \(k\)th individual after a random start? For bias questions, trace who is missing, who chooses whether to participate, and whether the way information is collected could influence responses. These distinctions build on Choosing the Best Sampling Method for a Situation and Identifying the Source of Bias in a Scenario.

Response routine: Identify the selection or response mechanism; name the sampling method or likely source of bias; explain which people may be overrepresented or left out; and state the population the results can reasonably describe.

A method’s name does not, by itself, prove that a sample is biased or unbiased. A random sample can still have undercoverage if its frame leaves out part of the population, and a survey can have nonresponse even when the initial selections were random. Conversely, a particular sample may happen to resemble the population even if the method was weak; that does not make the method a sound basis for generalizing.

Worked Example: Identify the Sampling Method

Worked Example: Inspecting Food-Delivery Orders

A fictional delivery company is reviewing orders placed during one month. For each plan, identify the sampling method and the detail that determines it.

Plan A: The company uses a random-number generator to select 80 individual orders from a complete list of all 2,400 orders from that month.

Solution: This is a simple random sample (SRS) of orders. The chance mechanism selects individual orders from the full frame, rather than first selecting groups or selecting from separate categories. The list covers the stated population of orders for that month, and every possible group of 80 orders has an equal chance of selection under an SRS plan.

Plan B: The company divides orders into four delivery zones, then randomly selects 20 orders from the list for each zone.

Solution: This is stratified random sampling. The zones are strata, and the plan randomly selects individuals from every stratum. The defining detail is that orders from all four zones are represented by selections made within each zone.

Plan C: The company randomly selects three delivery zones and reviews every order placed in those zones during the month.

Solution: This is cluster sampling. The plan randomly selects whole groups, or clusters, and includes every order within the selected clusters. It does not randomly select individual orders from every zone. This is the distinction practiced in Stratified Versus Cluster Sampling.

Plan D: The company picks a random starting order among the first 20 on an ordered list, then reviews every 20th order.

Solution: This is systematic random sampling: there is a random start followed by selection at a fixed interval. The company should also consider whether the ordering has a repeating pattern related to a feature of interest, as discussed in Systematic Random Sampling.

Conclusion: The number of orders reviewed does not identify the method. The selection mechanism does. Plan A selects individuals from the full frame; Plan B selects individuals from each stratum; Plan C selects whole clusters; and Plan D selects at a regular interval after a random start.

Worked Example: Diagnose Bias in a Poll

Worked Example: A Community Poll About Evening Bus Service

A fictional transit committee wants to learn whether adults in a town support adding evening bus service. It posts a poll on the transit department’s website and invites anyone to respond. Of the 1,200 responses, 870 support the proposal. A committee member says, “Most adults in town support the change.”

Identify the mechanism: The poll is a voluntary response sample. Adults decide for themselves whether to visit the website and submit an answer; the committee did not select a random sample of town adults.

Explain the concern: People with especially strong opinions about bus service may be more likely to seek out the poll and respond. If those with strong opinions are more likely to support or oppose the proposal than adults who do not respond, the results may overrepresent those views. The direction and size of any bias cannot be determined from the response count alone.

Check the frame and claim: The poll is available through the transit department’s website, not through a frame of all adults in town. Adults who do not use or visit the site may be less likely to see it. That creates an additional concern about undercoverage if the poll is treated as representing all town adults. The reported 870 of 1,200 responses describes the people who responded; it does not establish that most adults in town support the proposal.

Improve the plan: The committee could define the population as all adults living in the town, use a frame designed to cover that population, randomly select adults, and invite each selected person to answer the same fair question. It should track nonresponse and avoid assuming that nonrespondents hold the same views as respondents. A random selection improves the basis for generalizing, but the frame and response process still matter.

Worked Example: Design a Reproducible Plan

Worked Example: Surveying Public-Pool Members

A fictional city wants to estimate how many visits per month adult members make to its public pools during the summer. The membership roster lists 480 current adult members: 240 at North Pool, 144 at Central Pool, and 96 at South Pool. The city wants a sample of 48 members.

State: The population of interest is the 480 current adult members on the city’s roster. Each member is an observational unit. The variable is the number of visits that member makes to city public pools during the specified summer month. The sampling frame is the current roster, so the conclusion should be limited to members represented on that roster.

Plan: Use a stratified random sample, with pool membership location as the strata. This ensures that members from each pool are included. Allocate the 48 selections in proportion to the roster counts:

$$ 48\left(\frac{240}{480}\right)=24,\qquad 48\left(\frac{144}{480}\right)=14.4,\qquad 48\left(\frac{96}{480}\right)=9.6 $$

Because the allocations must be whole numbers and sum to 48, select 24 members from North Pool, 14 from Central Pool, and 10 from South Pool. Within each pool’s roster, assign unique labels. Use randInt(1, 240) to select 24 distinct North Pool members, randInt(1, 144) to select 14 distinct Central Pool members, and randInt(1, 96) to select 10 distinct South Pool members. For each list, accept a new label, skip repeats, and continue until the required number of distinct members has been selected. Match the labels to the roster and contact those selected members using the same survey question and procedure.

Conclude: This plan selects members by chance from all three strata and produces 24 + 14 + 10 = 48 members. The city can use responses to describe the members covered by the roster, while considering whether nonresponse might affect the results. The plan does not automatically represent nonmembers or anyone missing from the roster.

Worked Example: Separate a Random Draw From Nonresponse

Worked Example: A Survey of Repair Customers

A fictional appliance-repair business randomly selects 100 customers from its complete list of 800 customers served in the past year. It mails each selected customer the same survey about satisfaction. Sixty-five customers return it. A manager proposes selecting 35 additional customers from a list of people who recently visited the store to “make up” the difference.

Identify the original method: The initial selection is an SRS of 100 customers from the stated frame. The 65 returned surveys are the responses from those selected; the other 35 selected customers are nonrespondents.

Identify the proposed change: Adding people who recently visited the store is not a continuation of the original random selection plan. Those people were not selected from the original frame by the stated chance mechanism. They may differ from selected customers on satisfaction or other relevant characteristics, so treating them as replacements could introduce selection bias.

Improve the response plan: The business could send reminders or make follow-up contact with the originally selected nonrespondents, using the same survey. It should report that 65 of 100 selected customers responded, a response rate of \(65/100=0.65\), or 65%. Even with a random initial selection, nonresponse bias is possible if the nonrespondents’ satisfaction differs systematically from the respondents’ satisfaction.

Conclude: The 65 responses describe the respondents. The initial random selection supports a basis for generalizing to the 800 customers on the frame only if the coverage and response process are adequate. The manager should not imply that the extra store visitors were randomly selected members of the original sample.

How to Write a Strong Short Answer

Many sampling questions can be answered in a few precise sentences. The first sentence identifies the method or source of bias. The next sentence points to the relevant feature of the scenario. The final sentence explains the consequence for representation or the scope of the conclusion. This “name, evidence, consequence, scope” structure helps keep an answer tied to the situation instead of relying on a vague statement such as “the sample is unfair.”

For a plan-design question, connect the study question to a defined population, an appropriate frame, a selection method, and a clear variable or question. If you use a random method, state how the chance mechanism works. As in Describing a Random Sampling Plan in Context, another person should be able to follow the steps without inventing missing details.

Key takeaway: Identify a sampling method by what the procedure selects. Diagnose bias by tracing selection, coverage, response, and measurement. Then explain which group the resulting data can represent, and design plans with enough detail to reproduce the selection.

Common Mistakes and AP Exam Tips

  • Naming a method without the defining evidence. “It is cluster sampling” is stronger when followed by “because the plan randomly selects whole delivery zones and surveys every order in those zones.”
  • Calling every random method an SRS. Random selection within every stratum is stratified sampling; random selection of whole groups is cluster sampling. Describe what the chance process actually selects.
  • Confusing voluntary response with nonresponse. In voluntary response, people choose themselves into the sample. In nonresponse, selected individuals do not provide information. State which occurred in the scenario.
  • Calling any online survey automatically biased for one specific reason. Explain who could be missed, who chose to respond, or how the question or survey process could influence answers. A method name without a context-based mechanism is incomplete.
  • Claiming bias must push results in a particular direction. Give a direction only when the scenario supports it. Otherwise explain who may be overrepresented and why the direction cannot be determined from the information given.
  • Claiming that a large sample fixes a weak plan. A larger sample may reduce sampling variability under an appropriate method; it does not correct undercoverage, voluntary response, or nonresponse bias.
  • Generalizing beyond the frame. State the population the frame covers. Randomly sampling a frame cannot select people who are absent from it.
  • Writing a plan that cannot be carried out. State the frame, sample size, labels or selection mechanism, and rules for completing the selection. If the plan uses strata, explain how many individuals are selected from each.

On an AP response, use the context in every important claim. Say which people may be missed or more likely to respond, how that could affect representation, and what group the results describe. Avoid saying only “there is bias” or “the sample is random.” Explain the mechanism.

Check Your Understanding

Answer each item by naming the relevant method or concern and explaining the detail that supports your answer.

  1. A museum randomly selects 30 ticket holders from each of its four ticket-type lists. Identify the sampling method and explain what makes it that method.
  2. A neighborhood website asks residents to click a link and vote on a proposed park renovation. Identify a likely source of bias and explain which residents may be overrepresented.
  3. A recreation center randomly selects 50 members from its complete current roster, but only 28 return the survey. What is the initial sampling method, and what additional concern should the center consider?
  4. A town wants to survey 24 of the 240 current members of a community orchestra. Describe the frame, a chance-based selection procedure, and a stopping rule for an SRS.
  5. A company selects four warehouses at random and checks every package stored in those warehouses. Identify the sampling method and explain what is selected at random.