There Is No Single Best Sampling Method
In Sampling Variability in Repeated Samples, you saw that repeated random samples from the same population can produce different statistics. When planning a real sample, the selection method affects not only how practical the work is, but also which people can be represented and how much the results may vary.
The best method depends on the study question, the population, the sampling frame, and practical constraints such as travel, staff time, and access to records. A method that is efficient but misses part of the intended population may not answer the question well. Likewise, a statistically suitable method may be unrealistic if it requires reaching hundreds of distant locations.
As discussed in Population, Sampling Frame, and Sample and Why Random Selection Matters, a sampling frame is the list or source from which individuals are selected. Before deciding how to sample, check that the frame covers the population of interest. Random selection from an incomplete frame does not make the missing members eligible to be selected.
A Practical Decision Process
Start with the population and the kind of representation the study needs. Then ask what list or source is available and whether it contains the right individuals. The method should fit those facts, not just sound efficient in the abstract.
Name the individuals or units the question is about, along with relevant location and time boundaries.
Ask whether it covers the target population and whether selected individuals can be contacted or measured.
Consider whether important subgroups should each be represented, or whether individuals are naturally gathered in groups that are easier to reach.
Weigh time, travel, staff, cost, and possible patterns in the frame against the coverage and representation the question needs.
Identify a realistic concern, such as incomplete coverage, nonresponse, or similarity among members of selected clusters.
The earlier tutorials on Simple Random Sample Defined, Stratified Random Sampling, Cluster Sampling, and Systematic Random Sampling explain how each method works. The choice can be summarized like this:
| Method | Often useful when… | Practical question to check |
|---|---|---|
| Simple random sample (SRS) | A complete frame is available and every individual can be selected directly. | Can the selected individuals be reached without excessive time or travel? |
| Stratified random sample | Important subgroups should all be represented and group membership is known. | Can the population be divided into clear, nonoverlapping strata, with random selection within each? |
| Cluster sample | Individuals are naturally grouped, and reaching a few whole groups is much easier. | Are the selected groups chosen by chance, and might their members be unusually similar? |
| Systematic random sample | A usable ordered frame makes selecting every \(k\)th individual convenient. | Could a repeating pattern in the order line up with the interval? |
These are not guarantees that a method will work well in every setting. For example, stratification is helpful only if the strata are defined appropriately and the frame includes members of each group. A systematic plan is not safely random merely because it uses an interval: the start must be chosen at random from the first interval, and the frame order deserves attention.
Worked Example: Sampling Clinic Waiting Times
Worked Example: A Week of Scheduled Clinic Visits
A fictional clinic network wants to estimate the typical waiting time for patients with scheduled appointments who check in at one of its three clinics during a particular week. The check-in system records every such arrival and the time the patient is brought into the exam room. Staff can create a complete list of those arrivals after the week ends. The list contains 600 patients: 300 at Clinic A, 180 at Clinic B, and 120 at Clinic C. The network can measure 60 waiting times.
State: The population is all patients with scheduled appointments who check in at these three clinics during the specified week. The sampling frame is the completed check-in list of those arrivals, with their recorded times. It is not the appointment schedule, which could include people who did not arrive.
Plan: A stratified random sample by clinic is suitable if the network wants each clinic represented. The conditions to check are that the frame covers the stated population, each patient is assigned to exactly one clinic stratum, and staff can randomly select patients within each clinic. The method is feasible because the system can sort the arrival list by clinic. An SRS from the combined list would also be possible, but could happen to include relatively few patients from one clinic.
Do: Allocate the sample of 60 in proportion to the clinic counts. The population proportions are 300/600, 180/600, and 120/600. The corresponding sample sizes are:
Select 30 patients at random from Clinic A’s arrival list, 18 from Clinic B’s, and 12 from Clinic C’s. For example, staff could assign each arrival a unique label within its clinic and use a random-number generator to choose the required number of distinct labels. Every listed patient is eligible for selection, and every clinic contributes to the sample.
Conclude: This stratified plan fits the goal because it uses actual arrivals—not scheduled appointments alone—and guarantees representation from all three clinics. It supports conclusions about the stated population if the arrival records are complete and selected patients’ waiting times are available. If some arrivals are missing from the records, or some selected times are missing, those issues could limit how well the data represent all arrivals.
Worked Example: Reducing Travel With Cluster Sampling
Worked Example: Inspecting Household Water Filters
A fictional county wants to estimate whether households use their supplied water filters. Its 1,200 households are listed in 24 neighborhoods, with 50 households in each neighborhood. Interviewers can visit a few neighborhoods in person, but visiting households scattered across the entire county would require much more travel. The study can contact about 300 households.
State: The population is the county’s 1,200 households, and the frame lists all of them by neighborhood. Neighborhoods are natural clusters because each consists of a group of households located near one another.
Plan: Cluster sampling is a practical option if the main constraint is travel. Check that all neighborhoods are included in the frame, that the 24 neighborhoods form nonoverlapping groups covering the households, and that neighborhoods will be selected at random. A limitation is that households within a neighborhood could have similar filter-use habits, so a sample from a few neighborhoods may vary more than a sample spread across the county.
Do: Randomly select 6 of the 24 neighborhoods, then contact all 50 households in each selected neighborhood. The planned number of households is:
For comparison, an SRS of 300 households from the full list would also meet the target sample size, but those households might be scattered among many neighborhoods. That could increase travel time. The cluster plan instead concentrates the visits in 6 randomly selected locations.
Conclude: Cluster sampling is preferable here if reducing travel is important and the county accepts the possibility that households in the same neighborhood may be alike. Randomly selecting whole neighborhoods gives each neighborhood a chance to be included; selecting only the easiest neighborhoods would not be the same chance-based plan. The sample represents the county most defensibly if all neighborhoods are covered and the selected households respond.
Worked Example: Using an Ordered List Carefully
Worked Example: Reviewing Customer-Service Tickets
A fictional service center wants to sample 120 of the 8,400 tickets closed during one month. A complete electronic list is available, and tickets are ordered by close time. Staff can easily review every \(k\)th ticket. However, ticket types may follow a repeating schedule during the day, so staff need to consider whether that schedule could create a pattern in the ordered list.
State: The population is the 8,400 tickets closed during the month, and the frame is the complete ordered ticket list. The aim is to select 120 tickets for review.
Plan: Systematic random sampling would be convenient if the order does not have a repeating pattern that coincides with the sampling interval. Check that the frame includes all tickets in the target month, that each ticket has one position, and that the random start is selected from the first interval. If ticket types repeat at an interval related to \(k\), use an SRS or rearrange the frame before using a systematic plan.
Do: The interval is the frame size divided by the desired sample size:
Choose a random start from positions 1 through 70. Suppose the selected start is 23. Then select positions 23, 93, 163, and continue adding 70 until 120 tickets have been selected. The last position is \(23+119(70)=8{,}353\), which is within the 8,400-ticket frame.
Conclude: The systematic plan is efficient because it provides a clear, spread-out selection from the list with little repeated setup. Its suitability depends on checking the ordering: if a ticket-type schedule repeats in a way that aligns with every 70th position, the sample could overrepresent or underrepresent some types. In that case, a simple random sample from the complete list would be a safer choice.
When Practical Constraints Change the Choice
A method should not be chosen by convenience alone. Convenience sampling, discussed in Convenience Sampling and Its Bias, does not use chance to select individuals. For example, asking only the first 60 clinic patients who are available is not equivalent to randomly selecting 60 from the arrival list. The readily available patients could differ from other arrivals in ways related to waiting time.
Practical constraints can also change the population a study can reasonably describe. If a clinic has records only for patients who booked appointments, it cannot automatically use those records to represent every person who sought care, including walk-ins. The population, frame, and method must align. In the clinic example, the target was deliberately limited to scheduled patients who checked in, and the frame was built from their actual arrivals.
Nonresponse is another practical concern. A chance-based sample does not ensure that everyone selected will participate or have usable measurements. As in Nonresponse Bias, if selected individuals who do not respond differ from respondents on the variable of interest, the results may be biased. A realistic plan should allow for contacting selected individuals and should acknowledge the possibility of nonresponse rather than quietly replacing them with convenient alternatives.
The method also does not erase measurement problems. For instance, a complete list of arrivals will not guarantee accurate waiting times if timestamps are recorded inconsistently. As discussed in Sources of Variability in Collected Data, the variation in collected data can have more than one source. Method choice addresses selection, not every possible weakness in data collection.
Common Mistakes and AP Exam Tips
- Choosing a method before defining the population. First state who or what the question concerns and set the location and time boundaries. Then check whether the frame actually covers that group.
- Treating the sampling frame as the population. A list may omit some members or include people outside the target. Explain the mismatch and either improve the frame or narrow the population being described.
- Confusing stratified and cluster sampling. In stratified sampling, select individuals from every stratum. In cluster sampling, randomly select some clusters and include every member of those selected clusters.
- Claiming that a larger or easier sample is automatically better. A larger sample can reduce sampling variability under the same method, as discussed in Sample Size and Precision, but does not repair undercoverage or a nonrandom selection process.
- Ignoring patterns in a systematic frame. State the random-start rule and consider whether an ordering pattern could line up with the interval. If that risk is serious, recommend another method.
- Giving a method name without a reason. A full-credit justification connects the design to a specific feature of the situation—for example, known subgroups that must be represented, or travel costs reduced by sampling whole locations.
A concise justification can follow this pattern: name the target population and frame, state how chance selects the sample, explain why the method fits a practical or representational need, and identify a relevant limitation. Be specific: “stratify by clinic so each site is represented” is stronger than “stratified sampling is more accurate.”
Check Your Understanding
For each situation, choose a sampling method or identify information needed before choosing one. Justify your reasoning in context.
- A school has a complete list of students, with each student’s grade level recorded. The survey team wants students from all four grades represented. Which method is a sensible choice, and how would chance be used?
- A park district lists every household in a widely spread rural area. Staff want to interview residents but have limited travel funds. Explain when cluster sampling could be practical and name one limitation.
- A complete list of 2,400 event registrations is ordered by registration time. A team wants to select 80 registrations systematically. Find \(k\) and state two checks needed before using this method.
- A clinic wants to describe waiting times for patients who arrive without appointments, but its only list is the appointment schedule. Explain why selecting an SRS from that schedule does not directly meet the goal. What frame would better match the target?
- A team proposes interviewing people leaving a grocery store because they are easy to reach. Explain why this is not the same as a random sample of all county residents.