A Sampling Plan Should Be Reproducible
In Choosing the Best Sampling Method for a Situation, you compared methods based on the population, the available sampling frame, and practical needs. Now the task is to describe a plan precisely enough that another person could carry it out. Naming a method is not enough. A complete plan explains what list will be used, how individuals will be labeled, how chance will select labels, and what to do with the results.
This tutorial focuses on operational detail, especially for simple random samples (SRSs), as introduced in Simple Random Sample Defined and practiced in Selecting an SRS With randInt on a Calculator. The same care is useful when describing other random sampling methods. A good plan connects every step to the individuals in the study, rather than leaving the reader to guess how selection will happen.
The sampling frame is the list or source from which the sample is selected. As discussed in Population, Sampling Frame, and Sample, the frame should cover the population the study is intended to describe. A perfectly carried-out random draw from a list cannot select individuals who are missing from that list.
Four Details to State Clearly
For an SRS, the plan should let a reader answer four questions: Who is eligible? How are those individuals matched to labels? How will the random labels be generated? How will the selected labels become the sample? State the sample size, \(n\), so the reader also knows when to stop.
Describe who or what the study concerns and identify the specific roster or source that will be used. Say what makes a listed individual eligible.
Assign one label to each individual and make clear how to match each label back to that individual. No two people or units should share a label.
Name the method or tool and give its lower and upper bounds. For labels 1 through \(N\), for example, use randInt(1, \(N\)).
Process the generated labels in order. Include the individual for each new eligible label, skip a repeated label, and keep generating until \(n\) distinct individuals have been selected.
When using randInt, its endpoints are included: randInt(1, 80) can produce any whole number from 1 through 80. If a number appears again after that individual has already been selected, skip it. Continue generating numbers until the desired number of distinct labels has been accepted. This is sampling without replacement: an individual can enter the sample only once.
An existing ID number is not automatically a useful label. IDs may have gaps or fall outside the generator’s range. A clear plan can assign new consecutive labels to the eligible individuals and keep a roster that shows which person or unit corresponds to each label. If labels have leading zeros, explain that they identify roster positions; the random generator can still produce the corresponding whole numbers.
Worked Example: Selecting Rain Gauges
Worked Example: A Random Sample of Monitoring Sites
A fictional watershed team wants to check the calibration of 10 of its 80 rain gauges. The team has a complete roster of all 80 gauges currently operating in the watershed. Each gauge is an observational unit, and the team will record whether its measurement agrees with a reference reading.
State: The population is the 80 rain gauges currently operating in the watershed. The sampling frame is the team’s roster of those gauges. The intended sample size is \(n=10\).
Plan: Assign the gauges consecutive labels from 01 through 80, with each label matched to exactly one gauge on the roster. Use the calculator’s randInt function to generate integers from 1 through 80, inclusive. Read the output in order. For each label not already selected, add the matching gauge to the sample. Skip any label that has already been selected, and continue generating numbers until 10 distinct gauges have been chosen.
Do: Suppose the generator produces the sequence 17, 42, 17, 03, 68, 54, 29, 80, 11, 37, 06. The second 17 is a repeat, so it does not add a gauge to the sample. The first 10 distinct labels in order are 17, 42, 03, 68, 54, 29, 80, 11, 37, and 06. Match those labels to the roster; the sample consists of those 10 gauges.
Conclude: This is a complete SRS plan for the listed gauges because every gauge has one label, the random-number range matches the labels, and the rule produces 10 distinct selections. The plan supports describing the 80 operating gauges only to the extent that the roster includes them all. If the team’s goal were to describe gauges that are not on the roster, this sampling plan would not reach them.
Worked Example: Selecting Within Every Grade
Worked Example: A Stratified Plan for a Student Survey
A fictional high school wants to survey 36 students about the length of their travel time to school. Its complete roster contains 360 students: 150 in Grade 9, 120 in Grade 10, and 90 in Grade 11. The school wants each grade represented in proportion to its size.
State: The population is all 360 students on the school’s roster. The sampling frame records each student’s grade, so students can be separated into three nonoverlapping strata. The desired sample size is 36.
Plan: Use a stratified random sample, as described in Stratified Random Sampling. Allocate the 36 selections in proportion to the grade counts:
Within Grade 9, assign labels 1 through 150; within Grade 10, assign labels 1 through 120; within Grade 11, assign labels 1 through 90. Keep a separate roster for each grade so every label identifies exactly one student within its grade. Use randInt(1, 150) to select 15 distinct Grade 9 labels, randInt(1, 120) to select 12 distinct Grade 10 labels, and randInt(1, 90) to select 9 distinct Grade 11 labels. For each generator, skip repeats and continue until its required number of distinct labels has been selected. Match the labels to the students on the relevant grade roster.
Conclude: The plan selects students at random from every grade and produces 15 + 12 + 9 = 36 students altogether. The separate label ranges are intentional: the same number could identify one student in each of two different grades, but the grade-specific roster makes each selection unambiguous. The plan can represent students on the stated roster; it does not automatically represent students who are absent from it.
Worked Example: Making a Vague Plan Operational
Worked Example: Sampling Community Garden Plots
A fictional community garden has 125 active plots on a complete, current roster. The garden coordinator wants to inspect 8 plots to record whether each has a functioning water connection. A draft says, “We will randomly choose eight plots from the list.”
State: The population is the 125 active plots listed by the garden. The sampling frame is the complete current roster, and the sample size is \(n=8\).
Plan: The draft names a random choice but leaves out how it will be carried out. To make it reproducible, number the plots 1 through 125 and make a label-to-plot key. Use randInt(1, 125), with both endpoints included. Read the generated numbers in order, add the plot with each new label to the sample, skip any label that has already been selected, and continue until 8 distinct plots have been selected.
Do: Suppose the first generated numbers are 9, 41, 72, 41, 125, 1, 88, 53, and 27. The second 41 is skipped because plot 41 is already selected. The first 8 distinct labels are 9, 41, 72, 125, 1, 88, 53, and 27. The coordinator uses the key to find and inspect those 8 plots.
Conclude: The revised description tells another person which plots are eligible, how each label maps to a plot, what range to enter in the generator, how a repeat is handled, and when to stop. If one selected plot cannot be inspected, the coordinator should report that issue rather than silently substitute the nearest or easiest plot. A substitution based on convenience would change the selection procedure.
Write the Procedure, Not Just the Method Name
The examples show that even a correct method name does not fully describe a plan. “Take an SRS” does not tell a reader what the labels are, what values the generator can produce, or how to deal with a repeat. Similarly, “sample proportionally by grade” does not say how many students to select from each grade or how the random selection will happen within each one.
Match the bounds to the labeling scheme. If one roster has \(N\) individuals labeled 1 through \(N\), use randInt(1, \(N\)). If separate strata have their own rosters and labels, use the appropriate range for each stratum. Check that the sample size does not exceed the number of eligible individuals. For an SRS, the generator should select from the entire relevant frame, not just from a convenient portion of it.
The plan also needs to distinguish labels from selected people. A generated number is not yet a usable selection until it has been matched to an eligible individual on the roster. Keeping the roster and label key together helps prevent skipped numbers, misidentified individuals, or accidental duplicate selections.
Random selection and response are different stages. A selected person may decline or be unreachable. Do not describe replacing that person with someone nearby as part of the random draw; that would allow convenience to affect who enters the sample. As covered in Nonresponse Bias, nonresponse can limit how well the respondents represent the intended group. Describe how nonresponse will be recorded or addressed, and be cautious about what conclusions the completed responses support.
Common Mistakes and AP Exam Tips
- Writing only “choose randomly.” Give a specific chance mechanism, such as randInt, and the exact inclusive range it should use.
- Using unclear or repeated labels. Every eligible individual needs a unique label, and the plan should explain how to match each label to the roster.
- Forgetting the sample size or stopping rule. State how many distinct individuals are needed and say to continue until that many have been selected.
- Not explaining what happens to repeats. For selection without replacement, skip a label that has already been accepted and generate another number.
- Using a range that does not match the frame. If there are 125 eligible individuals labeled 1 through 125, a generator range of 1 through 100 could never select the last 25.
- Confusing a random plan with a guarantee of no bias. Random selection helps avoid choosing individuals based on preference or convenience, but it cannot fix an incomplete frame, nonresponse, or inaccurate measurement.
- Replacing an unavailable selection with an easy-to-reach person. That changes the chance-based procedure. State the nonresponse issue instead of treating a convenient substitute as if it had been randomly selected.
For a full-credit written response, a reader should be able to carry out the selection without inventing missing steps. A concise description can state the population and frame, the unique labeling scheme, the random generator and bounds, and the rule for repeats and stopping. Add the specific sample size and any important frame limitation.
Check Your Understanding
For each situation, identify what a complete random sampling plan should specify.
- A recreation center has a roster of 96 current members and wants an SRS of 12. What labels and randInt bounds could it use, and how should it handle repeated numbers?
- A library has 210 registered volunteers in three branches. A team wants 7 volunteers from each branch. What information should be labeled separately, and how should the generator range be chosen for each branch?
- A technician wants to inspect 15 of 180 solar panels. Explain why “use a random number generator” is incomplete unless the label range and selection rule are also given.
- A student selected by a random plan cannot be contacted. Why is choosing the nearest available student not part of the original random selection procedure?
- A roster of 75 people is numbered 1 through 75, but a proposed generator range is randInt(1, 60). Identify the problem with the proposed bounds.