What a Significant Result Can—and Cannot—Support
A news story may report a statistically significant result and then describe it as proof of a major breakthrough. The test result and the headline are separate claims: a small p-value can provide evidence against a null hypothesis, but it does not automatically establish a large effect, a cause, or a conclusion about everyone.
In “Common Misconceptions About Errors and Significance,” you learned that a significant result is evidence, not proof. Here, the task is to audit what is said after a test is reported. As in “Is a Statistically Significant Difference in Means Practically Important?”, the size of an estimated difference matters in its own right. Also check the study design: it determines whether a result can be generalized to a population or interpreted as evidence of a cause-and-effect relationship.
A useful new habit is to break a headline into distinct questions before judging it:
- Evidence: Does the reported p-value support rejecting \(H_0\) at the stated significance level?
- Size: How large is the estimated difference, in the response’s original units? Is it practically important in this setting?
- Reach: Does the study design support a conclusion about a population, or only about the people or units studied?
- Cause: Were treatments randomly assigned, or does the study show only an observed difference or association?
- Certainty: Does the report present evidence appropriately, or claim that the result proves a hypothesis?
These questions are related, but none answers all the others. A test may provide convincing evidence of a mean difference while that difference is small. A randomized experiment may support a causal conclusion for its participants without representing all people. A random sample may support a population conclusion without showing that one variable caused another.
Audit the Claim, Not Just the P-Value
First identify the population parameter and the hypotheses being tested. For a one-sample t test, the target may be a population mean \(\mu\); for paired data, it may be the population mean difference \(\mu_d\); and for independent groups, it may be the difference between two population means. The procedure should match the design, as discussed in “One-Sample t Versus Two-Sample t” and “Paired t Versus Two-Sample t.”
Then check the decision. If \(p\leq\alpha\), reject \(H_0\); if \(p>\alpha\), fail to reject \(H_0\). As “Writing Complete Conclusions for Mean Inference Problems” explains, a rejection can support wording such as “the data provide convincing evidence” for the alternative claim, in context. It does not support “the study proved” that claim.
Next, translate the estimate back into the original units. A significant result about a mean is not a statement that every individual’s response changed, nor does it tell us whether the change matters in practice. For a practical-importance judgment, consider the effect’s size and a relevant threshold; “Using Confidence Intervals to Judge Practical Importance” shows how an interval can add information about plausible effect sizes.
Finally, match the scope of the conclusion to how the data were collected. Appropriate random sampling can support generalization to the population sampled. Random assignment can support a cause-and-effect conclusion for the experiment’s participants. Neither feature automatically supplies the other. A voluntary survey, for example, does not become representative merely because its p-value is small.
Worked Example: A Lower Mean Is Not Proof the App Caused It
A fictional news report says, “A navigation app cuts commuters’ travel time.” The report describes a random sample of 25 people who use the app. Their mean commute time is 27.6 minutes, with a sample standard deviation of 4 minutes. For this example, the comparison benchmark is 30 minutes. The sample values are roughly symmetric with no apparent outliers, and the eligible app-user population is more than 250 people. Test whether the population mean commute time for app users is below 30 minutes at \(\alpha=0.05\), then assess the headline.
State. Let \(\mu\) be the population mean commute time, in minutes, for people in the population from which the app users were randomly sampled. The hypotheses are \(H_0:\mu=30\) and \(H_a:\mu<30\).
Plan. Use a one-sample t test for a population mean. The random sample condition is met by the stated sampling method. The 10% condition is met because 25 is less than 10% of a population larger than 250. The sample is small, so check its shape: the stated roughly symmetric pattern with no apparent outliers supports using the t procedure.
Do. The standard error is \(s/\sqrt{n}=4/\sqrt{25}=4/5=0.8\) minutes. The test statistic is
There are \(25-1=24\) degrees of freedom. A one-sided t calculation gives \(p\approx0.0031\), rounded. Since \(0.0031<0.05\), reject \(H_0\).
Conclude. The data provide convincing evidence that the population mean commute time for app users is less than 30 minutes. The headline’s causal verb, “cuts,” is not supported by this study: it samples app users but does not randomly assign app use or compare their outcomes under randomized app and no-app conditions. The result concerns a population mean, not every commuter’s travel time. A more careful report would say, “A random sample of app users had evidence of a mean commute time below the 30-minute benchmark.”
Significance Does Not Measure the Size of an Effect
A very small p-value can accompany a modest difference when the standard error is small. The test addresses how compatible the data are with \(H_0\) under the test model; it does not label the effect “large,” “dramatic,” or “important.” “Why Huge Samples Make Tiny Differences Significant” explains why a large sample can make a small difference statistically significant.
For a news claim about a treatment’s impact, look for the estimated difference in the response’s original units. If that estimate is 0.3 minutes, for example, “statistically significant” does not turn 0.3 minutes into a large reduction. Whether 0.3 minutes matters depends on the context and an appropriate practical benchmark. Also ask whether the design supports the report’s claim about people beyond those who took part.
Worked Example: A Significant Difference Can Still Be Small
In a fictional experiment, 400 participants are randomly assigned to each of two reminder systems. The mean time to complete a task is 8.0 minutes in the standard-system group and 7.7 minutes in the new-system group. Both groups have a sample standard deviation of 1.2 minutes. A report calls the new system “a dramatic time-saving breakthrough.” Test for a difference in population means at \(\alpha=0.05\), and evaluate the wording. For context, suppose the team considers a reduction of at least 2 minutes practically important.
State. Let \(\mu_N\) and \(\mu_S\) be the population mean completion times for the new and standard systems, respectively. Test \(H_0:\mu_N-\mu_S=0\) against \(H_a:\mu_N-\mu_S\ne0\).
Plan. Use an unpooled two-sample t test. The response is quantitative, and participants were randomly assigned to two independent groups. Each group contains 400 participants, so the sample sizes are large enough for the t procedure to be robust to departures from normality. The groups were created by random assignment; independence of participants’ responses is assumed from the design as described. Random assignment supports a causal comparison for these participants. It does not, by itself, establish that they represent all potential users.
Do. The estimated difference, new minus standard, is \(7.7-8.0=-0.30\) minutes. Its standard error is
The test statistic is \(t=(-0.30-0)/0.0849\approx-3.54\). With equal sample sizes and standard deviations, the unpooled degrees of freedom are \(400+400-2=798\). The two-sided p-value is approximately \(0.0004\), rounded. Since \(0.0004<0.05\), reject \(H_0\).
Conclude. The experiment provides convincing evidence of a difference in mean completion time, and random assignment supports the conclusion that the reminder system caused a difference for the participants. The estimated reduction is 0.30 minutes, far below the stated 2-minute practical-importance threshold. The result is statistically significant, but the word “dramatic” is not supported by the estimated size. The experiment also does not justify a broad claim about all potential users unless the participants were selected in a way that represents that population.
A Small P-Value Cannot Repair a Weak Design
Some reports focus on a test calculation while overlooking how participants entered the study. A small p-value does not remove selection bias, create a random sample, or turn an observational study into a randomized experiment. If a condition needed for the intended inference is not met, report what the calculation says under its assumptions, then explain why the intended population or causal conclusion is not justified.
Be precise about the limitation. A voluntary sample may still show a pattern among the people who responded. The concern is that those respondents may differ systematically from nonrespondents, so their result may not represent the target population. Similarly, a before-and-after change in the same people can be described for those participants, but without a comparison group or random assignment, other changes over time may help explain the observed difference.
Worked Example: A Significant Before-and-After Result From Volunteers
A fictional online story reports that a sleep tea “adds 18 minutes of sleep.” Sixty-four people who chose to respond to an online invitation recorded sleep duration before and after using the tea. Define each difference as after minus before. The mean difference is 18 minutes and the standard deviation of the differences is 40 minutes. A paired t test reports a two-sided p-value of about 0.0006 at \(\alpha=0.05\). Assess what the result supports.
State. If the target were a population of people using the tea, the relevant parameter would be \(\mu_d\), the true mean change in sleep duration, where each \(d\) is after minus before. The test’s hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\).
Plan. Paired measurements are appropriate because the same people recorded sleep before and after. However, the participants volunteered; they were not described as a random sample from a defined population. The 64 differences come from distinct respondents, but that alone does not satisfy the random sampling condition for generalizing to a population. The study also did not randomly assign tea use or include a comparison group, so a before-and-after difference cannot establish that the tea caused a change.
Do. The standard error of the sample mean difference is \(s_d/\sqrt{n}=40/\sqrt{64}=40/8=5\) minutes. The paired t statistic is
The degrees of freedom are \(64-1=63\). Under the paired t model, the two-sided p-value is approximately \(0.0006\), rounded. This calculation indicates that such a result would be unusual under a zero-mean-difference null if the model and conditions for the inference applied. The volunteers’ sampling method does not support a population inference, and the design does not establish causation.
Conclude. The 64 respondents had a mean increase of 18 minutes in their recorded sleep duration. The reported test calculation is statistically significant, but the voluntary sample does not justify concluding that the mean sleep change for a wider population differs from zero. The before-and-after design also does not show that the tea caused the change. A careful report would describe the observed result among respondents and state these design limitations rather than say the tea “adds” sleep.
Common Mistakes and AP Exam Tips
- Equating significance with a large effect: Report the estimated difference in original units before judging whether it is practically important. A p-value is not a measure of effect size.
- Turning a mean result into an individual claim: Evidence about a population mean does not mean every person’s response changed in the same direction or by the same amount.
- Generalizing from whoever responded: A voluntary sample is not automatically representative. State whether the data came from a suitable random sample before making a population claim.
- Using causal language for an observational comparison: Random sampling can support generalization, but causal conclusions require an appropriate randomized experiment. Describe an observed difference or association when the design does not establish cause.
- Writing “proved” after rejecting \(H_0\): A full-credit conclusion says the data provide convincing evidence for the alternative claim, in context. It does not claim certainty or say the test proved the alternative.
- Ignoring an unsupported headline while reporting a correct p-value: Evaluate the claim’s scope, effect size, and causal wording as well as the test decision. A correct calculation does not fix an overstatement.
On an AP response, separate the test conclusion from the headline’s extra claims. State the decision and evidence in context, describe the estimated effect in its units, and say whether the design supports generalization or causation. When one of those claims is unsupported, explain exactly which feature of the study limits it.
Check Your Understanding
For each question, distinguish the statistical test result from what the study design and estimated effect support.
- A news story calls a 0.2-minute mean reduction “a major improvement” after a significant test. What additional information is needed to judge the practical claim?
- A random sample of users has a mean outcome significantly below a fixed benchmark. Does that alone show that using the product caused the lower outcome? Explain.
- Why does random assignment support a causal conclusion for study participants but not automatically show that the result applies to everyone?
- A paired test of voluntary before-and-after survey responses is significant. What can be described about the respondents, and why might population and causal claims still be unjustified?
- Rewrite “The study proves the treatment works for all patients” as a careful conclusion when a randomized experiment found a significant mean difference among its participants, who were not randomly sampled.