Tutorials › AP Statistics › Practice Set: Formulating and Collecting Data

Statistical practices and exam synthesis · Tutorial 1017 of 1020

Practice Set: Formulating and Collecting Data

Work through a timed practice set that connects investigative questions to sampling and experimental design, then check how design choices shape the conclusions a study can support.

Intermediate 10 min read

What You'll Learn

  • Turn a broad topic into a focused question that names a population and measurable variables.
  • Choose and describe a probability sampling plan that fits a study’s goal.
  • Identify weaknesses in a sampling plan and explain how they could affect the results.
  • Plan an experiment using treatments, random assignment, replication, and control of other conditions.
  • Distinguish the conclusions supported by random sampling from those supported by random assignment.
  • Use a short timed set to practice concise, context-specific explanations.

Practice the First Two Statistical Practices

This set focuses on the work that comes before calculating statistics: deciding what to ask and how to collect data that can answer it. A strong question directs the design. The design, in turn, determines what conclusions are justified. These connections are central to “Four Statistical Practices in Context” and “Formulating Questions and Planning Data Collection.”

Set aside about 16 minutes for the three prompts below: roughly 5 minutes per prompt, followed by 1 minute to check your answers. Try each prompt before reading its worked solution. The suggested time is a guide, not a speed test. As in “Approaching a Multi-Part Free-Response Question,” identify what each part asks and answer it directly.

Key takeaway: A useful answer names the population or study subjects, explains how data will be collected, and matches the conclusion to the design. Do not let a broad topic or an attractive method replace a clear statistical question.

Timed Set: Question, Sample, Experiment

Read each scenario and respond to every part. For each question, note whether the goal is to describe a population, examine an association, or assess a possible cause. Then identify the data-collection method that fits that goal.

Worked Example: Turn a Broad Topic into an Investigative Question

Timed prompt. A public library is considering opening two hours earlier on Saturdays. The library director wants to know whether adult residents in the town support the change. State an investigative question that can be answered with data. Identify the population and a variable that would provide relevant data. In one sentence, explain why asking only current library visitors may not answer the director’s question well.

Step 1: Make the question specific. “Do people like the library?” is too broad: it does not say which people, what change they are judging, or what response will be measured. A focused investigative question is: What proportion of adult residents of the town support opening the public library two hours earlier on Saturdays? This is a descriptive question about a population proportion, not a question about whether earlier hours cause a change in library use.

Step 2: Name the population and variable. The population is all adult residents of the town. A relevant variable is each adult resident’s response to a clearly worded question about the proposed schedule, recorded as support or does not support. The wording and response choices should be neutral and understandable so that answers measure the view the director wants to learn about.

Step 3: Consider who is reached. Surveying only current library visitors could miss adults who rarely or never visit. Those residents might have different opinions or might become users if the hours changed. The visitors’ responses could therefore fail to represent all adult residents. This is a potential coverage problem, as discussed in “Identifying Bias and Confounding in Scenarios.”

Check the answer. The question specifies a population and a measurable response. A sampling plan should then give adults throughout that population a fair chance to be selected; the question alone does not make a convenience sample representative.

Worked Example: Choose a Sampling Plan for a School Survey

Timed prompt. A high school has 840 students: 210 in each of grades 9, 10, 11, and 12. The principal wants to estimate the proportion of all students who usually eat breakfast before school. The school can select 120 students from its enrollment roster. Describe a suitable sampling method, explain how the selection should be carried out, and identify one limitation that could remain even with this plan.

Step 1: Match the plan to the goal. The goal is to estimate a proportion for all enrolled students. Since each grade is a meaningful subgroup and has the same number of students, a stratified random sample can ensure representation from every grade. Divide the roster into four grade-level groups, then select a simple random sample of 30 students from each group. The total sample size is \(4(30)=120\).

Step 2: Make selection genuinely random. Assign each student in a grade a distinct number, then use a random method to select 30 numbers from that grade’s 210 students. Select without replacement and do not substitute students simply because someone is easier to contact. A fixed pattern, such as choosing every tenth name after sorting the roster by last name, would not automatically be random.

Step 3: Identify a remaining limitation. Some selected students may not respond, or they may report breakfast habits inaccurately. The random selection helps avoid choosing only convenient students, but it cannot guarantee responses or truthful, accurate answers. The school should make reasonable efforts to contact selected students and use neutral wording, while recognizing that nonresponse or response bias may still affect the results.

Why not just survey volunteers? A link sent to all students and answered by whoever chooses to respond would be a voluntary response sample. Students with especially strong opinions or interest in the topic might be more likely to reply. Random selection from the roster gives a more defensible basis for describing the student population.

Check the conclusion’s scope. If response is sufficiently complete and the survey is carried out as planned, the sample can support an estimate for students at this school. It does not, by itself, represent students at other schools. “Reading Study Descriptions for Design Flaws” and “Making Conclusions Consistent with Study Design” emphasize checking who could be represented before making a broader claim.

Worked Example: Design an Experiment About Study Reminders

Timed prompt. A school counselor wants to investigate whether a brief phone reminder improves students’ performance on a short study-skills quiz. The counselor can recruit 60 students selected at random from the school’s grade 10 roster. Describe an experiment comparing reminders with no reminders. Include the treatments, how students should be assigned, one way to control other conditions, and the conclusion the design could support.

Step 1: State the comparison. The explanatory variable is reminder condition, with two treatments: receiving a brief phone reminder before a scheduled study period, or receiving no phone reminder before that period. The response variable is the number of correct answers on the same short quiz, scored from 0 to 10. Defining the response precisely makes the comparison clear.

Step 2: Use random assignment. After the 60 students agree to participate, randomly assign 30 to each treatment. A random-number generator or another appropriate random method can make the assignments. Random assignment helps balance other student characteristics across the two groups, on average, so a difference in quiz performance can be attributed more plausibly to the reminder condition than it could in an observational study.

Step 3: Control other conditions and replicate. Have all students study for the same length of time, use the same study materials, and take the same quiz under the same conditions. Give the reminder group the planned message at the same point in the schedule, and do not give that message to the comparison group. With 30 students per treatment, the experiment has replication: each treatment is applied to multiple students, rather than a single student standing in for a condition.

Step 4: Set the limits on the conclusion. If the reminder group has a higher average quiz score, the experiment could provide evidence that the reminder caused higher quiz performance for these participants under these study conditions. Because the 60 participants were selected at random from the grade 10 roster, the result may also be generalized to that school’s grade 10 students, provided the selection and participation do not create important problems. It does not establish the effect for all students, other schools, or other kinds of reminders and quizzes.

Check the design logic. Random selection and random assignment do different jobs. Selection helps determine who the participants represent; assignment helps support a cause-and-effect conclusion. Having both is especially useful, but neither should be described as doing the other’s job.

How to Check Your Reasoning Under Time Pressure

After drafting an answer, make a quick design audit. First ask whether the question states what is being learned and about whom. Next ask whether the proposed data collection can answer that question. Finally, ask what the design permits the researcher to conclude. This brief audit can catch errors that a long explanation may not fix.

1
Identify the goal.
Is the question describing a population, examining an association, or investigating a possible cause?
2
Name the units and variables.
State who or what is observed and what information is recorded. Make the response or outcome specific enough to measure.
3
Describe selection or assignment.
For a survey, explain how individuals enter the sample. For an experiment, identify treatments and explain how subjects are assigned.
4
Limit the conclusion.
Use the study’s sampling and assignment methods to decide whom the results may represent and whether a causal claim is justified.

A clear response does not need to list every possible flaw. It should identify a relevant feature of the plan and explain its consequence. For example, “This sample is biased” is less informative than “Only current library visitors are asked, so adults who do not visit are left out and may have different views.” The second statement names who is missed and why that could matter.

Common Mistakes and AP Exam Tips

  • Writing a topic instead of a question. “Breakfast at school” does not specify what will be learned. Ask about a measurable feature, such as the proportion of students who eat breakfast before school.
  • Leaving the population unstated. “Most people support the change” is too broad when the study concerns adult residents of one town. Name the population that the plan is designed to represent.
  • Calling any large sample representative. A large voluntary response sample can still overrepresent people who choose to reply. Explain how the sample is selected; sample size alone does not remove selection bias.
  • Confusing random selection with random assignment. Random selection concerns choosing participants from a population. Random assignment concerns allocating experiment participants to treatments. State which one the scenario uses and what conclusion it helps support.
  • Claiming causation from a survey or observational study. An observed association does not establish that one variable caused another. A well-designed randomized experiment can support a causal conclusion for its study conditions.
  • Generalizing beyond the participants represented. Random assignment alone does not make a group representative of a larger population. Use the sampling method to set the population scope.
  • Listing an experimental feature without its purpose. Naming random assignment is not enough if the prompt asks why it matters. Explain that it helps create comparable treatment groups by balancing other factors, on average.

When time is short, prioritize the requested parts over a long opening paragraph. A direct answer that identifies the population, gives a workable selection or assignment method, and limits the conclusion is easier to evaluate than a broad explanation that leaves one of those tasks unfinished.

Key takeaway: Start with the investigative question, then choose a data-collection plan that can answer it. Random selection supports generalization to a population; random assignment supports cause-and-effect reasoning. Keep both conclusions within the design’s limits.

Check Your Understanding

Try these questions without looking back at the worked solutions. Focus on the link between the question, the data-collection plan, and the conclusion.

  1. A town wants to estimate the proportion of households that use a community compost service. State a focused investigative question and identify the population and one relevant variable.
  2. A school surveys the first 50 students who arrive at a morning event. Name one concern with this sampling plan and explain how it could affect the result.
  3. In the breakfast survey example, why is selecting 30 students at random from each grade a stratified random sample rather than a voluntary response sample?
  4. A researcher randomly assigns volunteers to two exercise routines but does not randomly select them from a population. What kind of conclusion can random assignment support, and what does it not establish by itself?
  5. Give one condition that should be kept the same for both groups in the study-reminder experiment, and explain why controlling it helps make the comparison useful.