Scan the Study Before Judging Its Conclusion
In “Matching Graphs and Summaries to Questions,” you matched evidence to the question being asked. Now apply that same question-first habit to the study description itself: before accepting a conclusion, check how the individuals entered the study and how any treatment was compared. A short summary may contain clues to important design limitations.
A study can have several weaknesses at once. A survey might miss part of its target population and receive replies from only some of the people contacted. An experiment might compare outcomes before and after a treatment without including a separate control group. The goal is not simply to name a flaw; it is to explain what that flaw makes difficult to conclude.
A useful distinction is between random sampling and random assignment. Random sampling concerns how individuals are selected from a population and can support generalizing results to that population. Random assignment concerns how participants in an experiment are placed into treatment groups and can support a cause-and-effect conclusion about the treatments. One does not automatically imply the other.
A Practical Design-Flaw Scan
Use the following questions as a checklist. A description may not provide enough information to answer every question. In that case, say what is not reported rather than assuming the study did something it never mentions.
Identify the target population and the individuals who actually had a chance to be included. Ask whether an important part of the target population was left out.
For a survey, check whether selected individuals responded. For an experiment, check whether those invited or eligible actually enrolled and completed the study.
If researchers evaluate a treatment, identify the group or condition used to show what would have happened without that treatment or with an alternative treatment.
Ask whether participants, researchers, or people measuring outcomes knew the treatment assignments. Consider whether expectations could affect behavior or measurement.
Determine whether participants were randomly assigned. If assignment was not random, consider whether the groups may have differed before treatment.
Explain which population claim, comparison, or cause-and-effect conclusion is weakened, and avoid claiming that a flaw proves the reported result is false.
Sampling and Participation: Who Is Missing?
Undercoverage occurs when some members of the target population are less likely or unable to be included in the sample because the sampling method does not adequately reach them. For example, a survey about all residents that uses only a website available to registered users may omit residents without access to that website. The issue is not just that the sample is small; it is that some parts of the population are missing or poorly represented in the way the sample is obtained.
Nonresponse occurs when individuals selected for a survey do not provide a response. Even if the initial selection process is appropriate, the people who reply may differ from those who do not. If people with a particular experience are more likely to answer, the responses may give a distorted picture of the target population. A low response rate can raise concern, but the rate alone does not tell us exactly how responses differ; the important question is whether responding is related to the variable being measured.
The two problems can occur together, but they are not interchangeable. Undercoverage concerns who could be reached or selected. Nonresponse concerns people who were selected or contacted but did not respond. A strong description names the specific group that may be missing and explains how its absence could affect the result.
Experiments: Look for Comparisons and Protection Against Bias
A control group provides a comparison for evaluating the treatment under study. Depending on the question, the control group might receive a placebo, receive the usual treatment, or receive no new treatment. The point is to compare outcomes under different conditions while keeping other aspects of the study as similar as possible.
When a study has no control group, a change observed after treatment may have other explanations. Participants might improve over time, change their behavior for reasons unrelated to the treatment, or be measured differently at the end. A before-and-after comparison can describe a change in the participants, but without a suitable comparison group it is difficult to attribute that change to the treatment alone.
Blinding means keeping participants or people involved in the study from knowing which treatment a participant received. When participants know, their expectations may affect their reported symptoms or behavior. When the people assessing an outcome know, their expectations may affect how they measure or record it. Blinding can reduce these risks, but it is not always practical or possible. Its absence is a potential concern to assess in context, not automatic proof that results are invalid.
Random assignment uses chance to place experimental participants into treatment groups. It helps create groups that are comparable at the start, on average, including with respect to factors researchers may not have measured. If participants choose their own groups or researchers assign groups in a systematic way, the groups may differ initially. Those pre-existing differences may help explain an observed outcome difference, so a cause-and-effect claim is less secure.
Worked Examples: Diagnose the Design and Its Limits
Worked Example: A Citywide Survey With Two Gaps
Situation. A fictional city wants to estimate how many adult residents use public transit each week. The survey team posts a questionnaire on the city’s recreation-center website and invites users to respond. The summary says 860 people completed it, but gives no information about how many were invited or how the respondents compare with other residents.
Scan the target and the route into the survey. The target population is all adult city residents, but the questionnaire is available through a recreation-center website. Residents who do not use that site may have little or no chance to take part. That creates a risk of undercoverage. If people who use the site differ in transit use from adults who do not, the respondents may not represent all adult residents.
There may also be a response issue: the invitation asks users to complete the questionnaire, so some people who see it may not reply. The summary does not state how many people saw or received the invitation and did not respond. Therefore, the extent of nonresponse cannot be evaluated from the information provided. Do not claim that a particular number refused or that nonresponse definitely changed the estimate.
State the limitation carefully. The 860 responses describe the people who completed this online questionnaire. Because the method may under-cover adults who do not use the recreation-center website, and because the response pattern is not reported, the result may not accurately estimate weekly transit use among all adult city residents. These limitations do not show that the estimate is necessarily too high or too low; the direction of any difference is not established.
What would improve the plan? The city could use a sampling method that gives adults across the city a chance to be selected, then make repeated contact attempts to encourage selected residents to respond. A more representative selection method addresses undercoverage; follow-up addresses nonresponse. Neither step guarantees that every response pattern will be free of bias, but each targets a different concern.
Worked Example: A Treatment Study Without a Control Group
Situation. In a fictional pilot, 32 adults with recurring wrist discomfort attend a four-week stretching program. Their average self-reported discomfort score decreases from 6.2 before the program to 4.8 afterward, on a scale from 0 to 10. The organizers conclude that the program reduced discomfort.
State the question. The claim is causal: it says the stretching program produced the decrease. The study includes a before-and-after comparison for participants, but no separate control group is described.
Plan the design check. Ask whether the observed change can be compared with what would have happened over those four weeks without the program or under another condition. Here, no such group is reported. The before-and-after scores show that the participants’ average reported discomfort was lower at the end, but other explanations for the change remain possible. Discomfort might change over time, participants might alter other activities, or their reports might be affected by expectations.
Interpret the numbers and the design. The observed average change is \(4.8-6.2=-1.4\) scale points, a decrease of 1.4 points among these participants. That is a description of the recorded before-and-after results, not by itself evidence that stretching caused the decrease. Because there is no control group for comparison, the study cannot separate the program’s effect from other explanations for change.
Conclude in context. The participants reported lower average discomfort after four weeks in the stretching program, but this study does not establish that the program caused the decrease. A comparison group observed over the same period would help assess whether the change is greater than changes that might occur without the program. Do not say the program had no effect; the design limitation means its causal effect is unclear.
Worked Example: Random Assignment but No Blinding
Situation. A fictional experiment randomly assigns 90 volunteers with mild seasonal symptoms to a new herbal drink or a similar-tasting drink without the active ingredient. Participants are told which drink they receive. At the end of two weeks, they rate symptom relief from 0 to 10. The researchers also know each participant’s assignment while collecting the ratings.
Identify the strengths first. There is a control condition, and participants are randomly assigned to the two drinks. Random assignment supports comparing the outcomes as evidence about the drinks’ effects on these volunteers, provided the study is carried out as described. The control group offers a comparison over the same period.
Identify the blinding concern. Participants know whether they received the new drink. Their expectations could affect how they experience symptoms or rate relief. The researchers collecting the ratings also know the assignments, which could affect how they ask questions or record responses. Because the outcome is self-reported and the assessment is not blinded, expectations could influence the measured results.
Keep the conclusion calibrated. If the assigned groups differ in reported relief, random assignment makes a cause-and-effect comparison more defensible for these volunteers than it would be in a non-randomized comparison. However, the lack of blinding creates a possible source of bias in the self-reported outcome. The design does not prove that expectations caused any difference, and the absence of blinding does not erase the random assignment. A careful conclusion reports the comparison and notes that unblinded reporting and assessment may affect its interpretation.
Consider an improvement. If feasible, the drinks could be made indistinguishable, and the people collecting or evaluating symptom ratings could be kept unaware of assignments. This would reduce the opportunity for expectations to influence reports or assessment. Blinding would not fix every possible limitation, but it would address the specific concern in this example.
Worked Example: A Non-Random Comparison of Two Programs
Situation. A fictional school compares attendance in two after-school tutoring programs. Students choose which program to attend. At the end of a term, students in Program A have higher average attendance than students in Program B. The coordinator says Program A improves attendance more.
Check how groups were formed. Students selected their own programs; the description does not report random assignment. The groups may differ before the programs begin. For instance, students who choose one program might have different schedules or initial attendance patterns from students who choose the other. These are plausible differences, not established facts about the students.
Distinguish comparison from causation. The reported averages describe the attendance of students in the two programs. They show an association between program membership and attendance in this comparison. Because group membership was not assigned at random, the difference may reflect pre-existing differences as well as any program effect. The summary does not allow us to isolate the cause of the higher average.
Conclude in context. Students who attended Program A had higher average attendance during the term than students who attended Program B. Because students chose their programs, this comparison does not establish that Program A caused higher attendance. Random assignment, if practical and appropriate, would help make the groups comparable at the start; otherwise, the conclusion should remain a cautious description of the observed groups.
Common Mistakes and AP Exam Tips
- Calling every sampling problem “nonresponse.” If some people could not be reached or selected in the first place, identify undercoverage. Reserve nonresponse for selected or contacted people who do not reply.
- Claiming a bias must point in a particular direction. A flaw may make an estimate less trustworthy, but the description may not establish whether the result is too high or too low. State the possible consequence without inventing its direction.
- Treating a before-and-after change as proof of a treatment effect. Without a comparison group, time-related changes and other explanations remain possible. Describe the observed change and limit the causal claim.
- Saying that no blinding makes a study useless. Lack of blinding is a potential concern, especially when outcomes are subjective or assessed by people who know assignments. Identify the pathway by which expectations could affect reports or measurements.
- Confusing random assignment with random sampling. Random assignment helps support cause-and-effect conclusions; random sampling helps support generalizing to a population. Name the feature the study actually used.
- Declaring that non-random assignment proves the treatment did nothing. It does not. It means groups may differ for reasons besides treatment, so a causal explanation is less secure.
- Listing a flaw without connecting it to the claim. A full-credit explanation identifies who may be missing, what comparison is absent, or how expectations may affect measurement, and then states the conclusion that is limited.
A useful response pattern is: name the feature, explain the risk, and limit the claim. For example: “Students chose their own tutoring programs, so the groups may have differed in attendance-related characteristics before the term. The observed attendance difference therefore does not establish that the program caused the difference.” This is more informative than writing only “there is bias.”
Check Your Understanding
For each description, identify the most relevant design concern and state what conclusion it limits.
- A county wants to survey all residents but mails questionnaires only to homeowners. What potential sampling issue should be considered, and why?
- A random sample of 500 people receives a survey, but only 110 respond. What is the concern, and what additional information would help assess it?
- Participants in a new sleep routine report their sleep quality before and after three weeks, with no other group observed. What can the results describe, and what causal conclusion is not established?
- In an experiment, the people rating photographs know which participants received the treatment. How could this lack of blinding affect the results?
- Two groups choose their own exercise plans, and one group later has a higher average fitness score. Why does this comparison not by itself establish a cause-and-effect relationship?