Tutorials › AP Statistics › Large Sample Sizes and Tiny P-Values

P-values and conclusions for proportions · Tutorial 492 of 1000

Large Sample Sizes and Tiny P-Values

Explore how sample size affects a one-proportion test’s p-value—and why a tiny p-value is not a measure of practical importance.

Intermediate 10 min read

What You'll Learn

  • Explain how increasing sample size changes the standard error and z statistic for a fixed difference from the null proportion.
  • Distinguish the p-value behavior for a two-sided test from that of a one-sided test.
  • Identify what happens when a fixed observed difference points opposite the one-sided alternative.
  • Interpret an extremely small p-value under the null model without treating it as the probability that the null hypothesis is true.
  • Separate statistical significance from the size or practical importance of a difference.

Why Sample Size Changes the Evidence

In P-Values for One-Sided Versus Two-Sided Tests, you learned that the alternative hypothesis determines which sample results count as at least as extreme as the observed result. Sample size also affects a p-value. When the observed difference from the null proportion is held fixed, a larger sample makes that difference larger relative to the null model’s standard error.

For a one-proportion \(z\)-test, the null standard error is \(SE_0=\sqrt{p_0(1-p_0)/n}\). Increasing \(n\) makes \(SE_0\) smaller. If the observed difference \(\hat{p}-p_0\) is fixed, the test statistic’s magnitude therefore grows. What that does to the p-value depends on the alternative: for a two-sided test, a larger absolute \(z\) gives a smaller p-value. For a one-sided test, the p-value gets smaller when the difference points in the direction of the alternative, but gets larger when the difference points in the opposite direction.

Key idea: More observations can make a fixed difference more statistically significant, but the direction of the alternative matters. In a one-sided test, a fixed difference opposite the alternative makes the p-value increase toward 1 as the sample size grows.

The formula shows the pattern. For a fixed difference \(d=\hat{p}-p_0\), the test statistic is \(z=d\sqrt{n}/\sqrt{p_0(1-p_0)}\). Its absolute value increases as \(n\) increases, unless \(d=0\). A two-sided test treats a more distant result in either direction as stronger evidence against \(H_0\). A one-sided test only treats results in its specified direction as stronger evidence against \(H_0\).

$$ z=\frac{\hat{p}-p_0}{\sqrt{p_0(1-p_0)/n}} =\frac{d\sqrt{n}}{\sqrt{p_0(1-p_0)}} $$

This is a comparison that holds the observed difference fixed; it does not mean that collecting more observations guarantees a particular sample proportion or a small p-value. Actual samples vary. Nor does it mean a tiny p-value proves that the difference matters in practice. As discussed in Statistical Significance Versus Practical Significance, the p-value measures how unusual the result is under the null model, not whether the effect is large enough to be useful.

Worked Examples: Sample Size and P-Values

Worked Example: A Fixed Difference in a Two-Sided Test

A fictional online grocery service wants to know whether the proportion of its account holders who choose reusable delivery bags differs from a benchmark of 40%. Consider two hypothetical random samples from a list of 200,000 account holders. In one sample, 42 of 100 account holders choose the bags. In the other, 4,200 of 10,000 do so. Both sample proportions are 42%, two percentage points above the benchmark. Test at \(\alpha=0.05\).

State: Let \(p\) be the true proportion of account holders on the service’s list who choose reusable delivery bags. For each sample, test \(H_0:p=0.40\) against \(H_a:p\ne0.40\). The question asks whether the proportion differs from 40%, so the alternative is two-sided.

Plan: Both samples are described as random samples, meeting the Random condition. For the sample of 100, \(100\leq0.10(200{,}000)=20{,}000\); for the sample of 10,000, \(10{,}000\leq20{,}000\). Thus, the 10% condition is met for each. Under \(H_0\), the expected success and failure counts are 40 and 60 for the first sample, and 4,000 and 6,000 for the second. All are at least 10, so the Large Counts condition is met for both one-proportion \(z\)-tests.

Do: In each sample, the observed proportion is \(\hat{p}=0.42\). For \(n=100\), the null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.40(0.60)}{100}} =\sqrt{0.0024} \approx0.04899,\qquad z=\frac{0.42-0.40}{0.04899} \approx0.4082 $$

The two-sided p-value is approximately \(2P(Z\geq0.4082)=0.6831\). Assuming the true proportion is 40%, a sample proportion at least this far from 40% in either direction would occur about 68.31% of the time under the test’s Normal model.

For \(n=10{,}000\), the calculation is:

$$ SE_0=\sqrt{\frac{0.40(0.60)}{10{,}000}} =\sqrt{0.000024} \approx0.004899,\qquad z=\frac{0.42-0.40}{0.004899} \approx4.0825 $$

The two-sided p-value is \(2P(Z\geq4.0825)\approx0.0000446\), rounded. Assuming the true proportion is 40%, a result at least this far from 40% in either direction would be very unusual under the model.

Conclude: For the sample of 100, \(0.6831>0.05\), so we fail to reject \(H_0\). It does not provide convincing evidence that the proportion of account holders who choose reusable bags differs from 40%. For the sample of 10,000, \(0.0000446\leq0.05\), so we reject \(H_0\). It provides convincing evidence that the proportion differs from 40%. The observed difference is two percentage points in both cases; the larger sample measures that difference more precisely, making it much more unusual under the null model.

Worked Example: A Fixed Difference Opposite the One-Sided Alternative

A fictional health clinic asks whether more than half of its registered patients would use a new online check-in option. To illustrate how sample size affects this one-sided test, compare random samples of 100 and 2,500 patients from a list of 50,000. In each sample, 46% say they would use the option. Use \(\alpha=0.05\).

State: Let \(p\) be the true proportion of registered patients on the list who would use online check-in. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\). The question asks whether the proportion is higher than half, so this is a right-tailed test. The observed proportion, 0.46, is below the null value and points opposite the alternative.

Plan: The samples are random, meeting the Random condition. The 10% condition is met because \(100\leq5{,}000\) and \(2{,}500\leq5{,}000\), where \(0.10(50{,}000)=5{,}000\). Under \(H_0\), the expected success and failure counts are 50 and 50 for \(n=100\), and 1,250 and 1,250 for \(n=2{,}500\). All counts are at least 10, so the Large Counts condition is met.

Do: For \(n=100\), \(\hat{p}=46/100=0.46\). The null standard error is \(\sqrt{0.50(0.50)/100}=0.05\), giving \(z=(0.46-0.50)/0.05=-0.80\). For the right-tailed test, the p-value is \(P(Z\geq-0.80)\approx0.7881\).

For \(n=2{,}500\), \(\hat{p}=1{,}150/2{,}500=0.46\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.50(0.50)}{2{,}500}} =\sqrt{0.0001} =0.01,\qquad z=\frac{0.46-0.50}{0.01} =-4.00 $$

The right-tailed p-value is \(P(Z\geq-4.00)\approx0.99997\), rounded. The observed result lies far in the direction opposite the right-tailed alternative, so almost all values in the null model’s distribution are at least as high as this observed \(z\).

Conclude: At \(\alpha=0.05\), both tests fail to reject \(H_0\). Neither provides convincing evidence that more than half of registered patients would use online check-in. In this comparison, the right-tailed p-value increases as \(n\) grows because the fixed difference points opposite the alternative. For comparison, a two-sided test would have a p-value of about 0.4237 for \(z=-0.80\), but about 0.0000633 for \(z=-4.00\). The test direction—not just the size of \(|z|\)—determines the one-sided p-value.

Worked Example: A Tiny P-Value for a Small Difference

A fictional packaging facility checks a random sample of 40,000 sealed cartons. The facility’s benchmark is that 30% of cartons have a particular recyclable label, but 12,400 in the sample have it. Test whether the true proportion differs from 30% at \(\alpha=0.05\).

State: Let \(p\) be the true proportion of cartons produced by this facility that have the recyclable label. Test \(H_0:p=0.30\) against \(H_a:p\ne0.30\).

Plan: The cartons were randomly sampled, meeting the Random condition. Suppose the facility produced at least 400,000 cartons in the relevant production period. Then \(40{,}000\leq0.10(400{,}000)=40{,}000\), so the 10% condition is met. Under \(H_0\), the expected counts are \(40{,}000(0.30)=12{,}000\) cartons with the label and \(40{,}000(0.70)=28{,}000\) without it. Both exceed 10, satisfying the Large Counts condition.

Do: The sample proportion is \(\hat{p}=12{,}400/40{,}000=0.31\), one percentage point above the null proportion. The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.30(0.70)}{40{,}000}} =\sqrt{0.00000525} \approx0.0022913,\qquad z=\frac{0.31-0.30}{0.0022913} \approx4.3644 $$

The two-sided p-value is \(2P(Z\geq4.3644)\approx0.0000127\), rounded. Assuming the true proportion is 30%, a sample proportion at least one percentage point away from 30% in either direction would be extremely unusual under the test’s Normal model.

Conclude: Since \(0.0000127\leq0.05\), we reject \(H_0\). The sample provides convincing evidence that the true proportion of cartons with the recyclable label differs from 30%. The p-value is tiny, but the observed difference is only one percentage point. Whether that difference is important for packaging decisions depends on practical considerations, not on the p-value alone.

What an Extreme P-Value Does—and Does Not—Say

A tiny p-value describes the rarity of the observed result, or a result at least as extreme, assuming the null hypothesis is true and the test’s conditions and model are appropriate. For example, the packaging result’s p-value of about 0.0000127 means that results this far from 30% in either direction would occur about 0.00127% of the time under that null model. It does not mean there is a 0.00127% chance that \(H_0\) is true.

Large samples can yield tiny p-values for small departures from a null value because their standard errors are smaller. That is useful when a small difference needs to be detected, but it also means statistical significance alone does not communicate the difference’s size. Report the sample proportion and its difference from \(p_0\), interpret the test result in context, and consider whether the difference matters for the real decision.

The direction qualification is essential. For a two-sided test, increasing the absolute test statistic makes the p-value smaller. For a one-sided test, increasing \(|z|\) makes the p-value smaller only when the observed result moves farther into the alternative’s direction. If it moves farther the other way, the one-sided p-value increases toward 1. The same \(z\) can therefore be compelling evidence in one direction and very weak evidence for an alternative pointing the other way.

Common Mistakes and AP Exam Tips

  • Saying a larger sample always makes the p-value smaller: That is the pattern for a two-sided test with a fixed nonzero difference, or for a one-sided test when the difference is in the alternative’s direction. State the direction qualification.
  • Using \(|z|\) alone to interpret a one-sided test: Check whether the sign of \(z\) matches \(H_a\). A large negative \(z\) is not strong evidence for \(H_a:p>p_0\); the right-tail p-value is close to 1.
  • Calling a tiny p-value the probability that \(H_0\) is true: A p-value is calculated under the assumption that \(H_0\) is true and describes sample results. It does not assign a probability to the hypothesis.
  • Equating statistical significance with importance: Name the estimated difference in context. A small difference can have a tiny p-value with a large sample, but practical importance requires context beyond the test.
  • Claiming sample size guarantees significance: Increasing \(n\) can make a fixed difference easier to detect, but actual sample proportions vary. A larger sample does not guarantee a particular result.
AP Exam Tip: State \(H_a\) before interpreting the test statistic. For a two-sided test, explain that a larger \(|z|\) leads to a smaller p-value. For a one-sided test, say whether the observed difference is in the alternative’s direction; only then describe whether increasing sample size makes the p-value smaller or larger. Interpret the p-value under \(H_0\), and give the decision and conclusion in context.

Key Takeaway

Increasing the sample size reduces the null standard error. With a fixed observed difference, that moves the test statistic farther from zero. For a two-sided test, this makes the p-value smaller; for a one-sided test, the p-value gets smaller only when the difference points in the direction of the alternative. A very small p-value indicates strong evidence against \(H_0\), not necessarily a large or practically important effect.

Key takeaway: Interpret sample size, direction, and effect size together. A tiny p-value says the data are unusual under the null model; it does not say the null is probably true or false, or that the observed difference is important in practice.

Check Your Understanding

Use the relationship between sample size, direction, and p-values to answer each question.

  1. In a two-sided test, the difference \(\hat{p}-p_0\) stays fixed and nonzero while \(n\) increases. What happens to the absolute \(z\) statistic and the p-value?
  2. A right-tailed test has a fixed negative difference \(\hat{p}-p_0\). As \(n\) increases, does its p-value decrease or increase? Explain using the direction of the alternative.
  3. Suppose a two-sided test produces a p-value of 0.0004. State what this means under the null model and one thing it does not mean.
  4. A very large sample produces a statistically significant difference of 0.3 percentage points. Why is it not enough to report only that the result is significant?
  5. For a one-sided test, what must you check before claiming that a larger \(|z|\) means a smaller p-value?