Tutorials › AP Statistics › Drawing Conclusions About Causation From Test Results

P-values and conclusions for proportions · Tutorial 496 of 1000

Drawing Conclusions About Causation From Test Results

See how random assignment changes what a significant test result can support—and why a small p-value alone does not establish causation.

Intermediate 9 min read

What You'll Learn

  • Distinguish random assignment from random sampling and identify the different role each plays.
  • Explain why random assignment can support causal conclusions when an experiment finds a statistically significant difference.
  • Use a two-proportion test result to make a cautious conclusion about a randomized experiment.
  • Explain why failing to reject a null hypothesis does not prove a treatment has no effect.
  • Recognize why a significant difference in an observational study supports an association, not a causal claim.
  • Write conclusions that match both the test result and how the data were collected.

When Does a Significant Result Support Causation?

A small p-value can provide convincing evidence against a null hypothesis, as discussed in Interpreting a P-Value in Context. But it does not, by itself, tell us whether one variable caused another. To make a causal conclusion, we need to consider how the data were collected—especially whether researchers randomly assigned individuals to treatments.

In a randomized experiment, researchers use chance to assign experimental units to treatment groups. If the experiment is carried out appropriately, a statistically significant difference in outcomes can support the conclusion that the assigned treatment caused a difference in the response for the experimental units studied. Random assignment helps make the groups comparable, on average, in both measured and unmeasured characteristics. It does not guarantee that the groups will be identical.

An observational study is different: researchers observe existing groups or choices rather than assigning treatments at random. A test may find a statistically significant association between group membership and an outcome, but other differences between the groups could help explain the result. Those differences are potential confounding variables. A small p-value does not remove them or turn an observational study into an experiment.

Definition: Random assignment uses chance to place experimental units into treatment groups. It can support cause-and-effect conclusions when an experiment finds convincing evidence of a difference. Random sampling uses chance to select individuals from a population; it supports generalizing results to that population, but it does not establish causation.

For a binary outcome, such as whether a participant meets a goal, we can compare the treatment and control group proportions. A two-proportion \(z\)-test evaluates whether the population proportions differ. In a test of equal proportions, the pooled sample proportion is used to calculate the standard error under the null hypothesis. The p-value then helps determine whether the observed difference is convincing evidence against equal proportions.

That statistical decision and the design-based conclusion are connected, but not interchangeable. A significant result from a well-conducted randomized experiment supports a causal claim about the assigned treatment and the participants studied. A significant result from an observational study supports an association, not a claim that one variable caused the other. And a result that is not statistically significant does not prove that a treatment has no effect.

Linking the Test to the Study Design

For a two-proportion test, let \(p_T\) be the true proportion of experimental units assigned to the treatment who would have the outcome of interest, and let \(p_C\) be the corresponding proportion for units assigned to the control. The null hypothesis of no difference is \(H_0:p_T=p_C\). The alternative might be \(H_a:p_T>p_C\) if the research question, set before examining the data, asks whether the treatment increases the outcome rate. A question about any difference instead uses \(H_a:p_T\ne p_C\).

A complete test still needs its conditions checked. For a randomized experiment, the random-assignment process supports the Random condition. If the study did not sample individuals without replacement from a finite population, the 10% condition does not apply to that experiment. Check the Large Counts condition using the pooled proportion under the null: each group should have at least 10 expected successes and 10 expected failures. Also consider whether the study was conducted so that one participant’s outcome did not affect another’s.

Conditions: For a two-proportion \(z\)-test, check that the data come from random assignment or appropriate random samples; check the 10% condition when sampling without replacement from a finite population; and verify that each group has at least 10 expected successes and 10 expected failures under the null. Consider whether the observations can reasonably be treated as independent.

The conditions support the test’s statistical calculation. They do not substitute for examining the design. In particular, a valid test for a difference in proportions cannot establish causation if the groups were formed by people’s choices or by an existing characteristic rather than by random assignment.

Worked Examples: What the Evidence Supports

Worked Example: A Significant Result in a Randomized Experiment

Researchers at a fictional school recruit 200 students who want to try a new study-planning reminder. By a random process, 100 are assigned to receive the reminder and 100 to a control group that does not receive it. At the end of the study period, 65 students in the reminder group and 45 in the control group meet a specified study goal. The research question, set in advance, is whether the reminder increases the proportion who meet the goal. Use \(\alpha=0.05\).

State: Let \(p_T\) be the true proportion of students assigned to receive the reminder who would meet the study goal, and let \(p_C\) be the corresponding proportion assigned to the control group. The hypotheses are \(H_0:p_T=p_C\) and \(H_a:p_T>p_C\). The alternative is right-tailed because the question asks whether the reminder increases the rate.

Plan: Students were randomly assigned to the two groups, so the Random condition is met. They were recruited for the experiment rather than sampled without replacement from a finite population, so the 10% condition is not applicable. The sample proportions are \(65/100=0.65\) and \(45/100=0.45\). Under \(H_0\), the pooled proportion is \((65+45)/(100+100)=110/200=0.55\). The expected success and failure counts in each group are \(100(0.55)=55\) and \(100(0.45)=45\), all at least 10. Assuming students’ outcomes do not influence one another, use a two-proportion \(z\)-test with a greater-than alternative.

Do: The observed treatment-minus-control difference is \(0.65-0.45=0.20\). The pooled standard error under the null is

$$ SE_0=\sqrt{0.55(0.45)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.00495} \approx 0.07036 $$

The test statistic is

$$ z=\frac{0.65-0.45}{0.07036} \approx 2.843 $$

For the right-tailed alternative, the p-value is approximately \(0.0022\), rounded. As a check, the one-tail area above \(z=2.843\) is about \(0.0022\). Since \(0.0022\leq0.05\), reject \(H_0\). If the two population proportions were equal, a treatment-minus-control difference of at least \(0.20\) would be unusual under the null model.

Conclude: The experiment provides convincing evidence that the reminder increases the proportion of participating students who meet the study goal. Because students were randomly assigned to the reminder or control group, this result supports a causal conclusion about the reminder for the students in this experiment. The study does not, on its own, establish that the same effect would occur for every student in other settings.

Worked Example: A Randomized Experiment Without a Significant Result

A fictional recreation program randomly assigns 50 participants to receive a new warm-up plan and 50 to follow the usual plan. After four weeks, 30 participants assigned to the new plan and 25 assigned to the usual plan meet a flexibility target. The researchers ask whether the new plan changes the proportion who meet the target, in either direction. Use \(\alpha=0.05\).

State: Let \(p_T\) be the true proportion of participants assigned to the new warm-up plan who meet the target, and \(p_C\) the corresponding proportion assigned to the usual plan. The hypotheses are \(H_0:p_T=p_C\) and \(H_a:p_T\ne p_C\).

Plan: Participants were randomly assigned, meeting the Random condition. The 10% condition is not applicable because the study assigns recruited participants to treatments rather than sampling without replacement from a finite population. The sample proportions are \(30/50=0.60\) and \(25/50=0.50\). The pooled proportion under the null is \((30+25)/(50+50)=55/100=0.55\). In each group, the expected numbers of successes and failures are \(50(0.55)=27.5\) and \(50(0.45)=22.5\), all at least 10. Assuming participants’ outcomes are independent, use a two-proportion \(z\)-test with a not-equal alternative.

Do: The observed difference is \(0.60-0.50=0.10\). The pooled standard error is

$$ SE_0=\sqrt{0.55(0.45)\left(\frac{1}{50}+\frac{1}{50}\right)} =\sqrt{0.0099} \approx 0.09950 $$

Thus,

$$ z=\frac{0.60-0.50}{0.09950} \approx 1.005 $$

For the two-sided alternative, the p-value is approximately \(0.3149\), rounded. Since \(0.3149>0.05\), fail to reject \(H_0\).

Conclude: The experiment does not provide convincing evidence that the new warm-up plan changes the proportion of participants who meet the flexibility target. Random assignment means that a convincing difference could have supported a causal conclusion, but this test does not provide that evidence. It does not prove that the plan has no effect.

Worked Example: A Significant Association in an Observational Study

A fictional school wants to learn whether students who regularly use an optional online tutoring platform are more likely to meet a math goal. Students choose whether to use the platform, so the researchers do not assign it. From a roster of 1,800 regular users, researchers randomly select 120; from a separate roster of 2,400 nonusers, they randomly select 120. Of the sampled users, 72 meet the goal; of the sampled nonusers, 48 meet it. Consider a two-sided test at \(\alpha=0.05\).

State: Let \(p_U\) be the true proportion of regular platform users in the defined school population who meet the math goal, and \(p_N\) the corresponding proportion of nonusers. The hypotheses are \(H_0:p_U=p_N\) and \(H_a:p_U\ne p_N\).

Plan: Students were randomly sampled within each group, meeting the Random condition for comparing these populations. The 10% condition holds because \(120\leq0.10(1{,}800)=180\) and \(120\leq0.10(2{,}400)=240\). The sample proportions are \(72/120=0.60\) and \(48/120=0.40\). The pooled proportion is \((72+48)/(120+120)=120/240=0.50\). Each group has 60 expected successes and 60 expected failures under the null, so the Large Counts condition is met. Use a two-proportion \(z\)-test. However, students chose whether to use the platform: there was no random assignment.

Do: The observed difference is \(0.60-0.40=0.20\). The pooled standard error and test statistic are

$$ SE_0=\sqrt{0.50(0.50)\left(\frac{1}{120}+\frac{1}{120}\right)} =\sqrt{0.0041667} \approx 0.06455 $$
$$ z=\frac{0.60-0.40}{0.06455} \approx 3.098 $$

The two-sided p-value is approximately \(0.0019\), rounded. Because \(0.0019\leq0.05\), reject \(H_0\).

Conclude: The data provide convincing evidence that the proportion meeting the math goal differs between regular users and nonusers in the defined school populations. This is evidence of an association, not that using the platform caused students to meet the goal. For example, students with stronger prior math skills might be more likely both to choose the platform and to meet the goal.

Why Random Assignment Matters

The randomized experiment in the first example gives a basis for comparing outcomes under two treatment assignments. Because assignment was determined by chance, pre-existing characteristics are not systematically assigned to one treatment group. Random assignment cannot guarantee that every characteristic is balanced in one particular experiment, but it gives a defensible basis for attributing a convincing difference in outcomes to the assigned treatment rather than to a consistent, pre-existing group difference.

In the observational example, students who chose the tutoring platform may differ from nonusers in ways related to meeting the goal. Prior skill, time available, or motivation are possible explanations. The test addresses whether the observed difference is surprising under equal population proportions; it does not determine why the proportions differ. A very small p-value makes the association difficult to explain by chance alone under the null model, but it does not rule out confounding.

Keep the scope of a causal conclusion tied to the experiment. Random assignment supports a cause-and-effect conclusion about the treatment assigned and the experimental units studied, provided the study was carried out appropriately. If participants do not follow their assigned treatments, outcomes are measured differently between groups, or many assigned participants drop out, interpretation may be more complicated. A test result cannot fix problems in how the experiment was conducted.

Also keep the statistical conclusion separate from a claim about importance. A significant difference can be small in practical terms, while a non-significant test can reflect limited information. As covered in Statistical Significance Versus Practical Significance, the p-value alone does not tell us whether a difference is large enough to matter.

Common Mistakes and AP Exam Tips

  • Claiming causation from a small p-value alone: A small p-value provides evidence against the null model. To support a causal conclusion, the study must use random assignment and be conducted appropriately.
  • Confusing random assignment with random sampling: Random assignment can support causation; random sampling can support generalization to a population. They answer different questions.
  • Calling an observational result causal: When people choose their groups or treatments, say the study found evidence of an association or difference. Do not say the exposure caused the outcome.
  • Writing “no effect” after failing to reject: The correct conclusion is that the data do not provide convincing evidence of a difference at the stated significance level. A non-significant result does not establish equal proportions.
  • Making the causal claim broader than the experiment: Identify the treatment, outcome, and experimental units in context. Do not assume that participants represent people or settings not included in the study.
  • Ignoring the study design after checking test conditions: Conditions justify using a statistical procedure; they do not establish that groups were randomly assigned. State how the groups were formed when explaining causation.
AP Exam Tip: Make two linked statements: first, state the test decision and what the data show about the proportions; then explain what the study design permits you to conclude. For a randomized experiment with a significant result, a cautious causal conclusion is appropriate. For an observational study, describe an association, not a cause.

Key Takeaway

A significant test result and a causal conclusion are not the same thing. Random assignment can make a causal interpretation reasonable when a well-conducted experiment finds convincing evidence of a difference. Without random assignment, even a very small p-value supports an association rather than proving that one variable caused another.

Key takeaway: Let the study design guide the conclusion. Random assignment plus convincing evidence can support a causal claim about the experimental units studied; a significant result from an observational study supports an association, not causation. Failing to reject a null hypothesis does not prove there is no effect.

Check Your Understanding

For each situation, distinguish what the test result says from what the data-collection method allows you to conclude.

  1. A randomized experiment finds a significant difference in the proportion of participants who reach a health goal. What kind of conclusion can the researchers make, and to whom should it be limited?
  2. A very small p-value is found in a study where people chose whether to use a study app. Why does that result not establish that the app caused a difference?
  3. In a well-conducted randomized experiment, a test fails to reject \(H_0:p_T=p_C\). Does this prove that the treatment has no effect? Explain.
  4. How do random assignment and random sampling differ in the conclusions they support?
  5. A significant test result shows an association between a behavior and an outcome in an observational study. Give one kind of factor that could help explain the association without the behavior causing the outcome.