Selecting Groups Instead of Individuals
In Stratified Random Sampling, we divide the population into groups and select individuals at random from every group. Another design begins with groups too, but uses them differently: we randomly select some entire groups and survey every member in each selected group. This is called cluster sampling.
For example, a city might want to ask residents about access to public parks. It could randomly select several apartment buildings and contact every household in those buildings. The buildings are the clusters. This may be more practical than making one citywide list of every household, especially if residents are spread across many locations.
Clusters must be defined so that every member of the population belongs to exactly one cluster. Together, the clusters must cover the population of interest. A cluster might be a classroom, a city block, a clinic, or a bus route, depending on the population and study question.
A cluster sample needs a sampling frame that lists the clusters and a way to identify all members within each selected cluster. As in Population, Sampling Frame, and Sample, a frame that leaves out part of the population or includes many people outside it can limit whom the results represent. Once clusters are selected, the plan must also make it possible to contact or measure every member in them.
How to Carry Out Cluster Sampling
First define the population and the groups that will serve as clusters. Then number the clusters and use a chance method to select some of them. Earlier tutorials on selecting an SRS with randInt or a random number table explain how to make that selection: the labels identify clusters rather than individual people. Finally, collect data from every member of each selected cluster.
Make sure the clusters do not overlap and together cover the population.
List the clusters and make sure the membership of each one can be identified.
Use a random method to choose the required clusters from the cluster frame.
Do not randomly choose only some members within a selected cluster if the plan is cluster sampling.
Report the selected clusters and the members surveyed, and note any nonresponse or missing coverage.
The final sample size can vary when clusters have different numbers of members. If the selected clusters contain \(m_1,m_2,\ldots,m_k\) members, then the sample includes all of those members:
Although clusters are selected randomly, the resulting sample is not generally an SRS of individuals. Members of the same selected cluster are included together, while members of clusters not selected are all left out. In an SRS, every possible group of a given size has an equal chance of being selected; cluster sampling does not generally have that property.
Worked Example: Selecting Community Centers
Worked Example: A Recreation Program Survey
A fictional city has 12 community centers and wants to ask participants about the hours of its recreation programs. The study will randomly select 3 centers and survey every current participant at each selected center. The center rosters contain 21, 18, and 24 participants at the selected centers.
State: The population is all current participants in the city’s recreation programs. The variable is each participant’s opinion about program hours. The community centers are the clusters.
Plan and check the design: Assume each participant is registered at exactly one center, so the center groups do not overlap, and all participants are listed on a center roster. The list of 12 centers is the cluster sampling frame. Label the centers 01 through 12, use a random method to select 3 distinct labels, and survey every participant on the rosters for those centers. Random selection must determine which centers are chosen; staff should not substitute a more convenient center.
Do: Suppose a chance method selects centers 04, 09, and 12. Survey all participants at each of those centers. The resulting sample size is:
The sample is 63 participants, not 3 participants: the 3 selected units at the selection stage are centers, and everyone in those centers is included in the sample.
Conclude: The study uses cluster sampling by randomly selecting 3 of the 12 community centers and surveying every participant at those centers. If the center frame covers the intended population and the selected participants respond, the random selection can support generalizing results to the population of participants represented by that frame. It does not justify conclusions about people who are not recreation-program participants.
Cluster Sampling Compared With Stratified Sampling
Both designs begin by dividing a population into groups, so it is easy to confuse them. The key difference is what happens after the groups are defined. Stratified random sampling selects individuals from every stratum. Cluster sampling selects some clusters and includes every member of those selected clusters.
| Feature | Stratified random sampling | Cluster sampling |
|---|---|---|
| What is selected? | Individuals from every stratum | Some whole clusters |
| Are all groups represented by design? | Yes; a sample is taken from every stratum | No; only selected clusters are included |
| What happens within a selected group? | Select some individuals at random | Include every member |
| Common reason to use the design | Ensure representation of important groups | Make data collection more practical across a spread-out population |
Imagine a school district wants information from students in every grade. It could stratify by grade and randomly sample students from each grade. That guarantees representation from all grades. Or it could select several classrooms at random and survey every student in those classrooms. That is cluster sampling, but the selected classrooms might not include every grade.
The distinction is about the sampling procedure, not just the names of the groups. Calling classrooms “strata” does not make a design stratified if the researcher selects only some classrooms and surveys everyone in them. Likewise, calling grade levels “clusters” does not make a design cluster sampling if the researcher samples individuals from every grade.
Worked Example: Choosing Between the Designs
Worked Example: A District Technology Survey
A fictional district has 1,200 students in 40 homerooms. A team wants to ask about access to a reliable internet connection at home. It is important to compare grades because the district wants to know whether reported access differs by grade.
Plan A—stratified random sampling: Use grade as the strata, then select students at random from every grade. This design ensures that all grades are represented. The team can choose sample sizes for the grades, including proportional allocation if it wants the sample’s grade proportions to match the district’s.
Plan B—cluster sampling: Number the 40 homerooms, randomly select some, and survey every student in those homerooms. This may reduce the work of visiting or coordinating with many classrooms. However, the selected homerooms might not include every grade, so this design does not guarantee grade representation for a grade-by-grade comparison.
Conclude: Because the district specifically wants to compare grades, stratified random sampling is a better match: it selects students from every grade. If coordinating a small number of classrooms is the main practical concern and the district does not require each grade to be represented, randomly selected homerooms could be a reasonable cluster design. The choice should follow the study’s goal, not just the fact that classrooms exist.
When Clusters Are Useful—and What They Can Cost
Cluster sampling can be practical when individuals are naturally organized into groups and contacting people one group at a time is easier than contacting them across the whole population. For example, a researcher might visit selected classrooms, selected city blocks, or selected clinics rather than travel to many scattered locations. Practicality does not make the sample automatically representative; the clusters still need to be selected by chance from an adequate frame.
The similarity of members within clusters also matters. Suppose a population is divided into many neighborhoods. If each neighborhood contains a broad mix of people with different views, and the neighborhoods are fairly similar to one another, selecting some whole neighborhoods may capture a range of views. But if neighborhoods differ greatly from one another and people within each neighborhood tend to have similar views, the particular neighborhoods selected can strongly affect the sample’s results.
This is a different pattern from the one that can make stratification useful. As discussed in Stratified Random Sampling, stratification can reduce sampling variability when members within each stratum tend to be similar on the measured variable and the strata differ from one another. For cluster sampling, selecting whole groups is often most informative when clusters are each reasonably varied and broadly resemble the population, rather than when each cluster is internally very similar and sharply different from other clusters.
This is a design principle, not a guarantee about a particular sample. A random selection can still happen to choose clusters with unusual results. If only a few clusters are selected, the sample may reflect differences among those particular groups more than differences among individuals throughout the population. Random selection reduces the opportunity for deliberate selection bias; it does not eliminate sampling variability, nonresponse, measurement problems, or frame problems.
Worked Example: Understanding the Selected Groups
Worked Example: Sampling Households on City Blocks
A fictional city has 30 residential blocks and wants to estimate the typical number of household recycling bins set out each week. Each block has a roster of households. The city randomly selects 5 blocks and counts bins for every household on those blocks.
Check the design: The population is the households in the city, and the blocks are clusters. For this design to cover the population, every household must be assigned to exactly one of the 30 blocks, and the block rosters together must cover the city’s households. The city must select the 5 blocks by chance and count bins for every household on those blocks.
Consider the group pattern: If recycling habits vary widely among households within each block and the blocks have similar mixes of households, each selected block may include a range of behavior. If some blocks have nearly universal recycling and others have very little, the selected blocks may have noticeably different averages. With only 5 blocks, which blocks happen to be selected could matter.
Conclude: The plan is cluster sampling if all households in the 5 randomly selected blocks are included. Whether it gives a stable estimate depends in part on how similar or different the blocks are on recycling behavior. The city should not claim that the sample necessarily represents all households perfectly just because the blocks were chosen randomly.
Common Mistakes and AP Exam Tips
- Sampling people within every group and calling it cluster sampling. That describes stratified random sampling when individuals are selected from every group. In cluster sampling, select some groups and include all their members.
- Selecting only some members of chosen clusters. If the plan randomly selects a subset of members within chosen clusters, it is not the cluster-sampling procedure defined here. State clearly whether every member of a selected cluster is included.
- Assuming every grade, neighborhood, or other group will be represented. Cluster sampling includes only the clusters selected by chance. Unlike stratified random sampling, it does not guarantee representation from every group.
- Using an incomplete cluster frame. Random selection from a list that omits clusters cannot represent the omitted part of the population. Check that the clusters cover the defined population.
- Claiming random selection removes all problems. Random selection helps limit selection bias, but nonresponse, inaccurate measurement, and undercoverage can still affect the results.
- Confusing the number of clusters with the sample size. Selecting 5 classrooms means 5 clusters were selected; if they contain 24, 26, 22, 25, and 23 students, the sample includes \(24+26+22+25+23=120\) students.
For a full-credit description, name the population and the clusters, explain that the clusters do not overlap and cover the population, describe how some clusters are selected at random, and state that every member of each selected cluster is included. To compare the design with stratification, say directly that stratified sampling selects individuals from every group, whereas cluster sampling selects some whole groups.
Check Your Understanding
Use the population, cluster-selection, and comparison ideas to answer each question.
- A district randomly selects 4 schools and surveys every student at those schools. Identify the clusters and explain why this is cluster sampling.
- A researcher selects a random sample of students from every school in a district. Is this cluster sampling or stratified random sampling? Explain the difference in the selection step.
- What must be true about a set of clusters for them to cover a population without overlap?
- A city selects a few whole neighborhoods. Explain why the sample may be affected if neighborhoods differ greatly in residents’ responses.
- A researcher selects 6 classrooms but surveys only 10 randomly chosen students in each. Does this match the cluster-sampling procedure in this tutorial? Explain.