A P-Value Does Not Speak for the Whole Study
A news summary may report a small p-value and then make a much broader claim: that a treatment works for everyone, that a result is almost certainly true, or that an observed difference is important. To evaluate such a summary, separate the statistical evidence from the study’s design and from the writer’s interpretation.
As you learned in Scope of Inference Based on How Data Were Collected, random sampling and random assignment support different conclusions. A p-value does not repair a weak sampling method, establish causation in an observational study, or make a small difference practically important. It summarizes how unusual the data, or more extreme data, would be if the null hypothesis and the test’s assumptions were true.
A useful audit asks four questions. What claim and population did the study actually test? Is the p-value described as a probability about data under the null, rather than a probability that a hypothesis is true? Does the study’s design support the stated scope and any causal language? Finally, does the report distinguish statistical evidence from practical importance?
Find the population, outcome, comparison, and null and alternative claims, if reported.
State what results would be at least as extreme under the null model and give the reported probability.
Ask how participants were selected and whether treatments or conditions were randomly assigned.
State the strength of evidence cautiously, then limit generalization and causal claims to what the design supports.
Worked Examples: Audit the News Summary
Worked Example: “A 98.4% Chance the Proposal Is Popular”
A fictional local news summary says: “A poll found that 62 of 100 randomly selected city adults support a proposal. A statistical test gives \(p=0.0164\), so there is a 98.4% chance that most city adults support it.” Assume the test compared the population proportion with \(0.50\) using a two-sided one-proportion \(z\)-test. The poll was drawn from a current city-adult registry with more than 1,000 people, and the test used \(\alpha=0.05\).
Identify the claim: Let \(p\) be the true proportion of adults represented by the city registry who support the proposal. The test evaluates \(H_0:p=0.50\) against \(H_a:p\ne0.50\). The sample proportion is \(\hat{p}=62/100=0.62\). Notice that the alternative asks whether support differs from one-half; it does not specifically test whether “most” adults support it.
Plan and check conditions: The poll used a random sample, so the Random condition is met as described. Since \(100\) is less than 10% of the registry population, the 10% condition is met. For the test’s Large Counts condition, use the null proportion: \(np_0=100(0.50)=50\) and \(n(1-p_0)=100(0.50)=50\); both are at least 10.
Do: The null standard error is \(\sqrt{0.50(0.50)/100}=0.05\). Therefore,
For the two-sided alternative, the p-value is \(2P(Z\ge2.40)\approx 0.0164\), rounded to four decimal places. Since \(0.0164<0.05\), reject \(H_0\).
Conclude and audit: The sample provides convincing evidence that the proportion of adults represented by this registry who support the proposal differs from 50%. The p-value means that, if the true proportion were 50% and the test’s conditions held, a sample result at least as far from 50% as 62 out of 100 would occur about 1.64% of the time. It does not mean there is a 98.4% chance that most adults support the proposal. The two-sided test alone does not establish that the proportion is above 50%; the sample’s observed proportion is above 50%, but the report should state the tested claim accurately. The random sample can support generalization to the population represented by the registry, subject to possible coverage and response limitations.
Worked Example: A Tiny P-Value from a Convenience Sample
A fictional technology blog reports: “At a gaming expo, 130 of 200 visitors preferred a new interface. A test against 50% gave \(p=0.0000221\), proving that gamers everywhere prefer it.” The visitors were recruited at booths and were not randomly selected. Assume the reported p-value came from a two-sided one-proportion \(z\)-test of \(H_0:p=0.50\).
Check the calculation and test setup: The observed proportion is \(\hat{p}=130/200=0.65\). Under the null, the standard error is \(\sqrt{0.50(0.50)/200}=\sqrt{0.00125}\approx0.03536\). Thus \(z=(0.65-0.50)/0.03536\approx4.243\). A two-sided Normal probability beyond \(|z|=4.243\) is approximately \(0.0000221\), rounded to seven decimal places. The null expected counts are \(200(0.50)=100\) successes and \(200(0.50)=100\) failures, so the Large Counts condition is met.
Evaluate the design: The sample is a convenience sample, not a random sample of gamers. The Random condition needed for the usual population inference is therefore not established. The calculation describes how unusual the observed result would be under the specified null model and its assumptions; the tiny reported value does not turn the expo visitors into a representative sample.
Conclude and audit: The sample shows that 65% of the recruited expo visitors preferred the interface. The test calculation indicates that this result would be very unusual under a 50% model, if the test model were appropriate. But the blog has not established that gamers everywhere—or even all expo visitors—prefer it. The strong phrase “proving” is also incorrect: a small p-value is evidence against the null model, not proof of a broad claim. The conclusion should be limited to the people observed unless a suitable sampling design supports a wider population.
Worked Example: An Association Reported as a Cause
A fictional health-news summary describes a survey of a random sample of adults from a county registry. Respondents reported their typical sleep duration and whether they had experienced a headache in the past week. The summary says that the test comparing headache proportions for adults sleeping fewer than six hours and those sleeping at least six hours gave \(p=0.04\). It then states: “Sleeping fewer than six hours causes headaches.” Assume the reported test was selected in advance and its conditions were judged adequate.
Interpret the evidence: The p-value of 0.04 means that, if the null hypothesis of equal population headache proportions were true, results at least as extreme as the observed difference would occur about 4% of the time under the test model. At a preselected \(\alpha=0.05\), the result is statistically significant and provides evidence of an association between reported sleep category and reported headaches in the population represented by the registry.
Check the scope: Random sampling can support generalizing the association to the population represented by the registry, provided the sampling frame and response process are adequate. However, respondents chose their sleep habits; researchers did not randomly assign people to sleep-duration groups. This is an observational study. Other factors could be related to both sleep and headaches, so the design does not support the claim that sleeping fewer than six hours caused headaches.
Improve the summary: A more defensible statement is: “In a random sample of county adults, the data provide evidence of an association between reported sleep duration and reported headaches.” The report should include the observed proportions or another measure of the difference so readers can judge its size. A p-value of 0.04 conveys evidence against equal proportions under the model; it does not by itself show that the difference is large, important, or causal.
Worked Example: Statistical Significance Is Not a Breakthrough
A fictional consumer report says: “In a random sample of 10,000 customers, 5,100 preferred a redesigned package. The test against equal preference gave \(p=0.0455\). The redesign is a major success.” Suppose the two-sided test compares the population preference proportion with \(0.50\), the sample was randomly selected from a customer list of more than 100,000, and \(\alpha=0.05\).
Check the test result: The sample proportion is \(\hat{p}=5100/10000=0.51\). The null standard error is \(\sqrt{0.50(0.50)/10000}=0.005\), so \(z=(0.51-0.50)/0.005=2.00\). The two-sided p-value is \(2P(Z\ge2.00)\approx0.0455\). Under the null, the expected success and failure counts are each \(10000(0.50)=5000\), satisfying the Large Counts condition. The random sample and 10% condition are met as described.
Evaluate the wording: Because \(0.0455<0.05\), the result is statistically significant at the stated level and provides evidence that the customer preference proportion differs from 50%. The observed difference is one percentage point. Whether that difference is a “major success” is a practical judgment that the p-value cannot settle. The report should distinguish the statistical result from a business decision and explain what size of preference shift would matter.
Conclude in scope: The random sample supports generalizing to the customers represented by the list, subject to its coverage and response limitations. It does not justify claims about all consumers or show that the redesign caused other outcomes, such as increased sales. A small p-value is not a measure of effect size or practical value.
Common Mistakes and What a Strong Audit Says
- Calling the p-value the chance that the null is true: The p-value is calculated assuming the null hypothesis is true. It is a probability about sample results under that assumption, not a probability that the null or alternative is true.
- Using “prove,” “guarantee,” or “certain”: A small p-value can provide convincing evidence against a null hypothesis when the method and conditions are appropriate. It does not prove a claim or eliminate uncertainty.
- Letting the p-value broaden the population: A p-value does not make a convenience sample representative. Name the population the sampling design can reasonably represent, as in the earlier tutorial on scope of inference.
- Claiming cause from an observed association: A small p-value does not remove confounding. Without random assignment in an appropriate experiment, describe an association or difference rather than a treatment effect.
- Confusing statistical significance with importance: A result can be statistically significant yet small in practical terms. Report the observed difference and discuss why that size matters in the setting.
- Repeating a headline without checking the alternative: A two-sided test of \(p=0.50\) against \(p\ne0.50\) tests for a difference in either direction. It does not, by itself, test the specific claim that the proportion exceeds 0.50.
- Ignoring what the summary leaves out: If the report does not give the sample size, the observed proportions, the population, or how the sample was selected, treat its conclusion cautiously and seek those details before endorsing its wording.
Key Takeaway
A news report can accurately quote a p-value and still draw an unsupported conclusion. Audit the tested claim and the p-value’s meaning, then use the sampling and assignment methods to set the limits of generalization and causation. Finally, separate evidence against a null hypothesis from the size or practical value of the observed difference.
Check Your Understanding
For each summary, identify the accurate interpretation and any unsupported wording.
- A random sample produces \(p=0.03\) for a two-sided test of whether a population proportion differs from 0.40. What does the p-value mean, and does it give the probability that the null hypothesis is true?
- A survey recruits volunteers at a music festival and reports a very small p-value before claiming that most people in the country share the volunteers’ preference. What design issue limits that claim?
- A randomized experiment finds a statistically significant difference between two treatments among volunteers. What kind of conclusion may random assignment support, and what does the volunteer recruitment limit?
- A large random sample finds a difference of less than one percentage point with \(p=0.02\). What additional information is needed before calling the difference important?
- A random sample survey finds an association between two self-reported behaviors, and a summary says one behavior causes the other. What feature of the design is missing for a causal conclusion?