Tutorials › AP Statistics › Mixed Practice: Planning a Statistical Study

Investigative questions and data collection · Tutorial 160 of 1000

Mixed Practice: Planning a Statistical Study

Learn to write a focused statistical question, match it to a workable study plan, and justify the scope of the conclusion the study could support.

Beginner 9 min read

What You'll Learn

  • Turn a broad prompt into a statistical question that names the group and outcome.
  • Match a study’s units, variables, and method to the question.
  • Use random selection and random assignment to justify the scope of a conclusion.
  • Distinguish conclusions about a population from conclusions about cause and effect.
  • Explain when a study supports conclusions only about the participants.

From Scenario to Question, Plan, and Scope

In Matching Questions to Data Collection Plans, we learned to connect a question to a population, variables, and a suitable method. An exam-style planning question often asks for one more step: explain what conclusions the proposed study could support. To do that, you must match the question and plan, then pay attention to how the study’s participants were selected and, if treatments are involved, assigned.

A strong answer is not just a list of study features. It is a chain of reasoning: the question names what we want to learn; the plan collects the information needed; and the conclusion stays within the evidence the plan can provide. A plan cannot establish that a treatment works before any data are collected. It can, however, make clear what a later result would allow researchers to conclude.

Key idea: Justify scope by identifying the group studied, whether participants were randomly selected, whether treatments were randomly assigned, and what was measured. Random selection can support generalizing to the population from which the sample was selected. Random assignment can support a cause-and-effect conclusion about the treatments. Neither one automatically guarantees the other.

Use the following exam-response sequence. It draws together ideas from Defining the Population of Interest, Deciding Which Variables to Measure, and Scope of Inference: Four Combinations. You do not need to repeat every definition from those tutorials; apply the ideas directly to the scenario.

1
State a statistical question.
Name the group and the outcome, comparison, or relationship to investigate. The question should anticipate variation across observational units.
2
Give a matching plan.
Identify the population, observational units, variables and their definitions, and the method for obtaining the data. Explain how participants or records will be selected when that matters.
3
Trace selection and assignment.
Say whether the plan uses random selection from the target population and whether it randomly assigns treatments. Do not treat these as interchangeable.
4
State the permitted scope.
Identify the population, if any, to which results could be generalized and whether a cause-and-effect conclusion is justified. If neither is supported, state that clearly.

When a scenario gives no results, do not invent a finding. Write what the study would be able to conclude if the data showed a particular pattern. When a scenario does report results, distinguish the study’s observed result from the broader conclusion that the design can support.

Worked Practice: A District Reading Survey

Worked Example: How Much Do Grade 8 Students Read?

A school district wants to describe how much time its Grade 8 students spend reading for pleasure. It has a current list of all Grade 8 students enrolled in the district. Give a statistical question, a plan, and a justified scope for conclusions.

Step 1: Statistical question. One suitable question is: “How many minutes did Grade 8 students enrolled in the district spend reading for pleasure during the past seven days?” Students may report different amounts, so the question anticipates variability. The time period and meaning of reading for pleasure are specific enough to guide measurement.

Step 2: Matching plan. The population is all Grade 8 students enrolled in the district during the current school year. The observational unit is one student. The variable is the student’s reported number of minutes spent reading for pleasure during the past seven days. The district could select a random sample from its current Grade 8 enrollment list and ask each selected student the same neutral survey question.

Step 3: Selection and assignment. The plan uses random selection from the district’s Grade 8 enrollment list. There is no treatment and no random assignment; students are reporting their existing behavior.

Step 4: Scope. If the random sample responds in a way that does not introduce substantial bias, the results could support a conclusion about Grade 8 students enrolled in this district during the specified school year. The study would describe reported reading time. It would not show that reading time causes any other outcome, because no treatment was assigned.

A concise exam response could say: “Select a random sample from the district’s current Grade 8 enrollment list and survey each selected student about the number of minutes they spent reading for pleasure in the past seven days. The observational unit is one student. Because the sample is randomly selected from the district’s Grade 8 students, the results could be generalized to that population, subject to possible nonresponse or inaccurate recall. Since this is a survey of existing behavior and no treatment is assigned, it cannot establish cause and effect.”

Notice the boundaries in the conclusion. The plan does not automatically represent students in other districts, students in other grades, or students enrolled in a different year. The question, sampling list, and conclusion refer to the same defined group.

Worked Practice: A Randomized Reminder Study

Worked Example: Do Assignment Reminders Increase On-Time Submissions?

A high school is considering an app reminder for Grade 10 students. The school can randomly select 120 students from its current Grade 10 enrollment list. The selected students can then be randomly assigned to receive either an app reminder before each of four weekly assignments or the school’s usual instructions. Give a question, plan, and justified scope.

Step 1: Statistical question. Ask: “For Grade 10 students at this high school during the current term, does receiving an app reminder before each weekly assignment increase the average number of the four assignments submitted on time?” This names the comparison, target group, and outcome. The outcome is a count from zero to four for each student.

Step 2: Matching plan. Randomly select 120 students from the school’s current Grade 10 enrollment list. The observational units are the selected students. Randomly assign 60 students to receive the reminders and 60 to the usual-instructions condition. For each student, record the number of the four assignments submitted by the stated deadlines. Use the same assignments and deadlines for both groups.

Step 3: Selection and assignment. The plan includes both random selection from the school’s Grade 10 students and random assignment to the two conditions. Selection and assignment serve different purposes: selection relates to generalizing beyond the students in the study, while assignment relates to comparing the effects of the two conditions.

Step 4: Scope. If the study is carried out as planned, with the assigned conditions maintained and outcomes recorded consistently, a difference between the groups can support a cause-and-effect conclusion about receiving the reminders rather than usual instructions. Because the students were randomly selected from this school’s Grade 10 enrollment, the results could also be generalized to that population for the current term, subject to limitations such as nonparticipation or students sharing reminders. The study does not automatically support conclusions about other schools or grade levels.

Since no results are provided, the correct response does not say that reminders increase submissions. It explains what a result could establish: “If the reminder group has a greater average number of on-time submissions, the randomized experiment would provide evidence that the reminders increased on-time submissions for the Grade 10 students represented by this study. The random sample supports generalizing the result to the school’s current Grade 10 population, assuming the sample and study are carried out appropriately.”

This is a case where both kinds of randomization matter. If the students were randomly assigned but recruited only as volunteers, the design could still support a cause-and-effect conclusion for the participants, but generalizing to all Grade 10 students would be less well supported. If the students were randomly selected but assigned to groups based on their existing app use, the sample could support generalizing a described association, but that comparison would not establish that the app caused a difference.

Worked Practice: An Observational Study With Limited Scope

Worked Example: Screen Time and Sleep in a Club

A student club wants to investigate whether daily recreational screen time is related to sleep duration among its members. The club plans to ask members who attend one meeting to complete a voluntary questionnaire. Give a statistical question, improve the measurement plan, and explain the scope.

Step 1: Statistical question. Ask: “Among members of the school’s student technology club, how is reported daily recreational screen time related to reported sleep duration on school nights?” The question names the group and asks about a relationship between two variables.

Step 2: Matching plan. The population of interest is all members of the student technology club during the current school year. One member is an observational unit. The questionnaire should define recreational screen time as the member’s estimated hours per day using a phone, tablet, computer, or television for recreation during the previous seven days. It should define school-night sleep duration as the member’s estimated hours of sleep on a typical night before a school day. Ask both questions of each participating member so that the two responses can be paired.

Step 3: Selection and assignment. Members are not randomly selected: only those attending the meeting who choose to respond are included. No treatment is assigned. The plan is an observational survey of existing behavior.

Step 4: Scope. The data could describe an association between the two reported variables among the respondents. Because the respondents are a voluntary group at one meeting, the study does not provide a strong basis for generalizing to all club members, let alone all students at the school. Because screen time is not assigned, an observed relationship would not establish that screen time causes a change in sleep duration.

A useful limitation statement is specific: “The results describe the responding members at the meeting. They may not represent members who did not attend or chose not to respond, and the observational survey cannot establish that screen time causes differences in sleep.” Simply saying “the study is biased” is less informative because it does not identify the selection issue or its consequence.

The question and variables can be well defined even when the sampling method limits the conclusion. A sound measurement plan cannot, by itself, repair a lack of random selection or create random assignment. As in Evaluating a Study Description Critically, assess the actual collection process rather than relying on a study’s label.

Common Mistakes and Full-Credit Communication

  • Writing a topic instead of a question. “Study reading” does not say what will be learned. Ask a question that names the group and the outcome, comparison, or relationship.
  • Leaving the variable vague. “Measure sleep” could refer to hours, quality, or a broad opinion. Define what is recorded and specify a time period or units when appropriate.
  • Claiming generalization without checking selection. Random assignment to treatments does not make participants a random sample of a broader population. Name the population from which random selection actually occurred.
  • Claiming causation without random assignment. A survey or observational study can reveal an association, but it does not establish that changing one variable caused a change in another.
  • Making the conclusion broader than the plan. A sample from one school does not automatically represent students across a region. Match the conclusion to the population represented by the selection process.
  • Reporting a result that the scenario never gives. If no data or outcome is reported, describe what the study could conclude if a specified pattern appeared. Do not predict or invent the study’s findings.

For full credit, connect each claim to a design feature. “The study can generalize” needs the population and a reason, such as random selection from that population. “The treatment caused a difference” needs a treatment comparison and random assignment. If a condition is missing, say what is supported instead: a description or association for the people who provided data.

Key takeaway: For a planning scenario, write a statistical question, propose a plan that measures the needed variables, and justify the scope using the actual selection and assignment methods. Keep the conclusion within the population and type of claim the design can support.

Check Your Understanding

For each scenario, identify a suitable statistical question, outline a matching plan, and state the scope of conclusion the plan could support.

  1. A town wants to estimate how many minutes residents spend walking for exercise in a typical week. It has a current list of residents and can select a random sample. What should the survey measure, and to whom could results apply?
  2. A teacher randomly assigns students in one class to use one of two review formats, then compares their next quiz scores. What can random assignment support, and what does it not establish about students in other classes?
  3. A student asks only friends who are present at lunch whether they bring reusable containers. What group do the responses describe, and what limits generalization?
  4. A school observes that students who report more sleep also tend to have higher quiz scores. Why can the observation support an association but not, by itself, a cause-and-effect conclusion?
  5. A proposed experiment randomly selects students from a school roster and then randomly assigns them to a new study app or usual study resources. Name one suitable response variable and explain how selection and assignment contribute differently to the study’s scope.