Tutorials › AP Statistics › Small Samples and Large P-Values

P-values and conclusions for proportions · Tutorial 493 of 1000

Small Samples and Large P-Values

Explore how limited data can leave a real difference undetected by a significance test, and learn what a large p-value does—and does not—tell you.

Intermediate 9 min read

What You'll Learn

  • Explain how a real difference in a population can coexist with a large p-value from a small sample.
  • Calculate and interpret two-sided p-values for one-proportion z-tests in context.
  • Compare tests with the same observed difference but different sample sizes.
  • Check the Random, 10%, and Large Counts conditions for each test.
  • Distinguish failing to reject the null hypothesis from evidence that the null hypothesis is true.

When a Real Difference Is Hard to Detect

In Large Sample Sizes and Tiny P-Values, you saw that increasing the sample size can make a fixed difference from the null proportion more statistically significant. The other side of that idea matters just as much: when a sample is small, a genuine difference in the population can produce a large p-value.

A test does not know the true population proportion. It compares the observed sample result with what would be expected if the null hypothesis were true. With limited data, ordinary sample-to-sample variation can be large relative to the difference being studied. As a result, the observed difference may not be unusual enough under the null model to provide convincing evidence against it.

Key idea: A large p-value means the observed result is not especially unusual under the null model. It does not establish that the null hypothesis is true, nor does it rule out a real difference in the population.

For a one-proportion \(z\)-test, the null standard error is \(SE_0=\sqrt{p_0(1-p_0)/n}\). A smaller \(n\) generally gives a larger standard error. So the same difference between \(\hat{p}\) and \(p_0\) produces a test statistic closer to zero, and typically a larger p-value in a two-sided test. The p-value is calculated assuming \(H_0\) is true; the alternative hypothesis determines which results count as at least as extreme as the observed result.

$$ z=\frac{\hat{p}-p_0}{\sqrt{p_0(1-p_0)/n}} $$

The examples below stipulate a true population proportion that differs from the null value. That lets us see how a real difference could be present even when a test does not detect it convincingly. In an actual investigation, the true proportion is unknown; a large p-value cannot tell us whether the population proportion is exactly the null value or differs from it.

Worked Examples: Limited Data and Large P-Values

Worked Example: A Real Ten-Percentage-Point Difference in a Small Sample

Imagine that a fictional community program sends a reminder to residents about a local service. For illustration, suppose the true proportion of residents who would respond after receiving the reminder is 0.60. A researcher, however, tests whether the population proportion differs from a benchmark of 0.50 using a random sample of 25 residents from a list of 500. Fifteen respond. Use \(\alpha=0.05\).

State: Let \(p\) be the true proportion of residents on the program’s list who would respond after receiving the reminder. Test \(H_0:p=0.50\) against \(H_a:p\ne0.50\). The question asks whether the proportion differs from 0.50, so the test is two-sided. Although the example stipulates that the true proportion is 0.60, the researcher does not know that from the sample.

Plan: The residents were randomly sampled, so the Random condition is met. The 10% condition is met because \(25\leq0.10(500)=50\). Under \(H_0\), the expected numbers of responses and nonresponses are \(25(0.50)=12.5\) and \(25(0.50)=12.5\). Both are at least 10, so the Large Counts condition is met.

Do: The sample proportion is \(\hat{p}=15/25=0.60\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.50(0.50)}{25}} =\sqrt{0.01} =0.10,\qquad z=\frac{0.60-0.50}{0.10}=1.00 $$

The two-sided p-value is \(2P(Z\geq1.00)=2(0.1587)=0.3173\), rounded. Assuming the true proportion is 0.50, a sample proportion at least 0.10 away from 0.50 in either direction would occur about 31.73% of the time under the test’s Normal model.

Conclude: Since \(0.3173>0.05\), we fail to reject \(H_0\). The sample does not provide convincing evidence that the proportion of residents who would respond differs from 0.50. But that is not evidence that the proportion is 0.50: in the stipulated scenario, it is actually 0.60. The small sample has produced a result that is compatible with the null model, even though a real difference exists.

Worked Example: The Same Difference with More Data

Keep the same fictional setting and the same stipulated true proportion of 0.60. Now suppose a separate random sample of 100 residents is taken from a list of 10,000, and 60 respond. Test \(H_0:p=0.50\) against \(H_a:p\ne0.50\) at \(\alpha=0.05\). The observed difference from the null proportion is still 0.10.

State: Let \(p\) again be the true proportion of residents on the program’s list who would respond after receiving the reminder. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: The sample is random, meeting the Random condition. The 10% condition is met because \(100\leq0.10(10{,}000)=1{,}000\). Under \(H_0\), the expected numbers of responses and nonresponses are both \(100(0.50)=50\). Both are at least 10, so the Large Counts condition is met.

Do: Here, \(\hat{p}=60/100=0.60\), the same sample proportion as in the first example. The calculations are:

$$ SE_0=\sqrt{\frac{0.50(0.50)}{100}} =\sqrt{0.0025} =0.05,\qquad z=\frac{0.60-0.50}{0.05}=2.00 $$

The two-sided p-value is \(2P(Z\geq2.00)\approx0.0455\). Assuming the true proportion is 0.50, a sample proportion at least 0.10 away from 0.50 in either direction would occur about 4.55% of the time under the test’s Normal model.

Conclude: Since \(0.0455\leq0.05\), we reject \(H_0\). This sample provides convincing evidence that the proportion of residents who would respond differs from 0.50. The observed difference is the same ten percentage points as in the first example, but the larger sample gives a smaller null standard error and a smaller p-value. This comparison illustrates why limited data can fail to reveal a real effect convincingly.

Worked Example: A Large Observed Difference That Is Still Not Significant

A fictional environmental group studies whether residents in a town sort food scraps for composting. The group wants to know whether the proportion differs from a benchmark of 0.20. It takes a random sample of 50 residents from a list of 2,500; 15 say they sort food scraps. For illustration, suppose the true population proportion is 0.30. Test at \(\alpha=0.05\).

State: Let \(p\) be the true proportion of residents on the town list who sort food scraps for composting. Test \(H_0:p=0.20\) against \(H_a:p\ne0.20\). The sample proportion is 0.30, ten percentage points above the null value, but the direction of the test is two-sided because the question asks whether the proportion differs.

Plan: The residents were randomly sampled, meeting the Random condition. The 10% condition is met because \(50\leq0.10(2{,}500)=250\). Under \(H_0\), the expected number who sort food scraps is \(50(0.20)=10\), and the expected number who do not is \(50(0.80)=40\). Both counts are at least 10, so the Large Counts condition is met.

Do: The sample proportion is \(\hat{p}=15/50=0.30\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.20(0.80)}{50}} =\sqrt{0.0032} \approx0.05657,\qquad z=\frac{0.30-0.20}{0.05657} \approx1.7678 $$

The two-sided p-value is \(2P(Z\geq1.7678)\approx0.0771\), rounded. Assuming the true proportion is 0.20, a sample proportion at least ten percentage points away from 0.20 in either direction would occur about 7.71% of the time under the test’s Normal model.

Conclude: Since \(0.0771>0.05\), we fail to reject \(H_0\). The data do not provide convincing evidence that the proportion of residents who sort food scraps differs from 0.20. The sample difference is substantial in the context of the example, and the stipulated true proportion is 0.30, but this sample does not cross the chosen significance threshold. A difference can be real and noticeable while the available evidence is not strong enough for rejection.

Why a Large P-Value Does Not Support the Null Hypothesis

The phrase “fail to reject” is deliberately limited. It reports that the sample did not provide enough evidence, under the test’s decision rule, to reject \(H_0\). It does not say that the sample has verified \(H_0\), that \(H_0\) is probably true, or that the alternative is false. As discussed in Why We Never Accept the Null Hypothesis, a test’s failure to reject is not an acceptance of the null.

A large p-value says something precise about the data and the null model: assuming \(H_0\) is true, results at least as extreme as the observed one are not rare according to the test. It is not a probability assigned to \(H_0\). In particular, it is not the probability that the null is true or the probability that a real effect does not exist.

Limited data can make the standard error large, so a difference that would be unusual in a larger sample may not be unusual in a smaller one. The first two examples held the observed difference fixed and showed how sample size changed the test result. In real sampling, the observed sample proportion varies from sample to sample as well, so taking a larger sample does not guarantee a small p-value. It can, however, make a fixed difference easier to distinguish from the null value.

When interpreting a large p-value, report the observed proportion and the result of the test, but do not turn the result into a claim that there is no effect. If the size of the difference matters to a real decision, the context and practical consequences matter too. A significance test alone does not settle those questions.

Common Mistakes and AP Exam Tips

  • Writing “the null hypothesis is true” after failing to reject: Say that the data do not provide convincing evidence for the alternative. Failing to reject is not proof of \(H_0\).
  • Calling the p-value the chance that the null is true: State that the p-value is calculated assuming \(H_0\) is true and describes the chance of results at least as extreme as the observed result. The alternative determines which results count as at least as extreme.
  • Assuming a real difference must produce a small p-value: Sample results vary. With limited data, a real population difference can yield a result that is not unusual under the null model.
  • Using “large p-value” to mean “no effect”: A large p-value means the test did not find convincing evidence against \(H_0\). It does not establish that the population proportion equals \(p_0\).
  • Ignoring the observed difference: A test decision does not erase the sample result. Give \(\hat{p}\) and describe its difference from \(p_0\) in context, while keeping the conclusion cautious.
AP Exam Tip: A complete conclusion compares the p-value with \(\alpha\), states “fail to reject \(H_0\)” when the p-value is greater than \(\alpha\), and explains in context that the data do not provide convincing evidence for the alternative. Do not write that the null has been proved or accepted.

Key Takeaway

A small sample can produce a large p-value even when the population proportion truly differs from the null value. Limited data can leave substantial sample-to-sample variation relative to the observed difference. The correct conclusion is about the strength of evidence in the data, not a declaration that the null hypothesis is true.

Key takeaway: A large p-value means the observed result is not unusual under the null model; it does not support the null as true or rule out a real effect. When evidence is limited, fail to reject \(H_0\) and describe the conclusion cautiously in context.

Check Your Understanding

Use the relationship between sample size, observed differences, and evidence to answer each question.

  1. A test has \(p_0=0.50\), \(n=25\), and \(\hat{p}=0.60\). What is the null standard error, and why might the p-value be large despite the ten-percentage-point difference?
  2. In the first two worked examples, why did the same observed sample proportion lead to different p-values?
  3. A researcher fails to reject \(H_0\) in a test with a large p-value. Write one correct conclusion and one conclusion the researcher should not make.
  4. Explain what the p-value describes, including the role of the null hypothesis and the alternative hypothesis.
  5. Can a large p-value prove that there is no real difference in the population? Explain why or why not.