Tutorials › AP Statistics › Statistical Significance Versus Practical Significance

One-proportion hypothesis tests · Tutorial 477 of 1000

Statistical Significance Versus Practical Significance

See how sample size can make a small difference statistically significant, and why practical importance requires context beyond the p-value.

Intermediate 10 min read

What You'll Learn

  • Distinguish statistical significance from practical significance in a one-proportion test.
  • Explain how a large sample can make a tiny difference statistically significant.
  • Interpret the difference between a sample proportion and its null benchmark in context.
  • Connect a proportion difference to its scale and a prespecified practical threshold.
  • Avoid claiming that a small p-value proves an effect is important or that a large p-value proves no effect.

When “Significant” Does Not Mean “Important”

A hypothesis test can identify a difference that is unlikely to be explained by chance variation under a null model. But “statistically significant” does not automatically mean that the difference is large enough to matter in a decision. With a very large sample, even a tiny difference can produce a small p-value.

As explained in The Logic of a Significance Test, a p-value describes how unusual the sample result would be if the null hypothesis were true. It does not measure the size or importance of a difference. To discuss practical significance, also consider the size of the estimated difference, the consequences of that difference, the number of people or cases affected, and any practical threshold set for the situation.

Definition: A result is statistically significant at significance level \(\alpha\) when the p-value is less than or equal to \(\alpha\), leading to rejection of \(H_0\). Practical significance concerns whether the size and consequences of a difference are meaningful in the context of a decision.

For a one-proportion test, the observed difference from the null benchmark is \(\hat p-p_0\). Its size is expressed in proportion units; multiplying by 100 expresses it in percentage points. For example, a change from 50.0% to 50.1% is a difference of 0.1 percentage points, not 0.1 percent of the original value.

Statistical significance and practical significance answer different questions. A test against \(H_0:p=p_0\) asks whether the data provide convincing evidence that the population proportion differs from the benchmark. It does not ask whether the difference is large enough to meet a practical threshold. That judgment depends on the context, and a sample estimate alone does not establish the exact population difference.

Why Sample Size Matters

In a one-proportion \(z\)-test, the null standard error is \(SE_0=\sqrt{p_0(1-p_0)/n}\). As the sample size \(n\) increases, this standard error decreases. Consequently, a fixed difference between \(\hat p\) and \(p_0\) can produce a larger absolute \(z\) statistic and a smaller p-value in a larger sample.

That is why a very small difference may be statistically significant when \(n\) is large. The test is detecting evidence that the proportion is not exactly equal to the benchmark; it is not declaring the difference consequential. Conversely, a potentially important observed difference in a smaller sample may not be statistically significant because the sample result is too uncertain to provide convincing evidence.

Key distinction: The p-value reflects how compatible the data are with the null model, considering sample size and sample variation. The observed difference describes the estimated size of the effect. Practical importance requires interpreting that size in context.

A practical threshold should come from the situation, not be chosen after seeing which interpretation is convenient. It might be a minimum change worth the cost of a program, a maximum acceptable failure rate, or a number of affected cases that would alter a policy. When no threshold is given, explain what information would be needed to judge practical importance rather than inventing a universal cutoff.

Worked Examples

Worked Example: A Tiny Difference in a Million Accounts

A service team takes a random sample of 1,000,000 accounts from 15,000,000 eligible accounts. In the sample, 501,000 account holders use a particular feature. The historical benchmark is 50%. Test whether the population proportion differs from 0.50 at \(\alpha=0.05\). For this scenario, the team has said in advance that a difference smaller than 1 percentage point would not, by itself, justify redesigning the feature.

State: Let \(p\) be the true proportion of eligible account holders who use the feature. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\). A two-sided alternative is appropriate because the question asks whether the proportion differs from the benchmark in either direction.

Plan: Use a one-proportion \(z\)-test. The data come from a random sample, so the Random condition is met for inference to the eligible accounts represented by the sampling frame. Because sampling is without replacement, check the 10% condition: \(1{,}000{,}000\leq0.10(15{,}000{,}000)=1{,}500{,}000\), so it is met. Under the null, the expected success count is \(np_0=1{,}000{,}000(0.50)=500{,}000\), and the expected failure count is \(n(1-p_0)=1{,}000{,}000(0.50)=500{,}000\). Both are at least 10, so the Large Counts condition is met.

Do: The sample proportion and null standard error are:

$$ \hat p=\frac{x}{n} =\frac{501{,}000}{1{,}000{,}000} =0.501 $$
$$ SE_0=\sqrt{\frac{p_0(1-p_0)}{n}} =\sqrt{\frac{0.50(0.50)}{1{,}000{,}000}} =0.0005 $$

The test statistic and two-sided p-value are:

$$ z=\frac{\hat p-p_0}{SE_0} =\frac{0.501-0.50}{0.0005} =2.00 $$

The p-value is \(2\text{normalcdf}(2.00,1\text{E}99,0,1)\approx0.0455\), rounded. Since \(0.0455\leq0.05\), reject \(H_0\).

Conclude: The sample provides convincing evidence that the true feature-use proportion among eligible account holders differs from 50%. The observed proportion is higher: its difference from the benchmark is \(0.501-0.50=0.001\), or 0.1 percentage points. That observed difference is below the team’s 1-percentage-point redesign threshold, so the result is statistically significant but the sample estimate alone does not show a difference large enough to meet that practical threshold. The test was against exact equality, not against the practical threshold. At a scale of 1,000,000 comparable accounts, a difference of 0.001 corresponds to about 1,000 accounts, which could still matter depending on the costs and consequences. Neither the p-value nor the sample estimate settles that decision on its own.

Worked Example: A Potentially Meaningful Difference That Is Not Significant

A district randomly samples 1,000 families from 40,000 eligible families to ask whether they used a new online scheduling tool. Of those sampled, 510 report using it. The district compares the result with a 50% benchmark. A difference of about 1 percentage point would be large enough to merit further investigation. Test at \(\alpha=0.05\).

State: Let \(p\) be the true proportion of eligible families who used the tool. Use \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: Use a one-proportion \(z\)-test. The random sample supports the Random condition for the eligible families represented by the sampling frame. The 10% condition is met because \(1{,}000\leq0.10(40{,}000)=4{,}000\). Under \(H_0\), the expected success and failure counts are each \(1{,}000(0.50)=500\), so both are at least 10 and the Large Counts condition is met.

Do: The sample proportion is \(\hat p=510/1{,}000=0.51\). The null standard error and test statistic are:

$$ SE_0=\sqrt{\frac{0.50(0.50)}{1{,}000}} \approx0.01581 \qquad z=\frac{0.51-0.50}{0.01581} \approx0.6325 $$

For the two-sided alternative, the p-value is \(2\text{normalcdf}(0.6325,1\text{E}99,0,1)\approx0.5271\), rounded. Since \(0.5271>0.05\), fail to reject \(H_0\).

Conclude: These data do not provide convincing evidence that the proportion of eligible families using the tool differs from 50%. The observed difference is 1 percentage point, which the district considers worth investigating, but this sample does not establish that the population proportion differs from the benchmark. Failing to reject \(H_0\) does not prove that the true difference is zero or that the tool has no practical value. The district could consider collecting more data or examining costs and outcomes before making a decision.

Worked Example: A Small Difference That Matters at Large Scale

A fictional subscription service randomly samples 250,000 of its 8,000,000 eligible customers. In the sample, 125,750 renewed. The service compares its renewal proportion with a 50% benchmark. Its planning team has determined that, at a scale of 5,000,000 comparable customers, about 10,000 additional renewals would be operationally important. Test for a difference from 50% at \(\alpha=0.05\), then discuss the practical context.

State: Let \(p\) be the true proportion of eligible customers who renew. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\).

Plan: Use a one-proportion \(z\)-test. The random sample meets the Random condition for the eligible customers represented by the sampling frame. The 10% condition is met because \(250{,}000\leq0.10(8{,}000{,}000)=800{,}000\). Under \(H_0\), the expected success count is \(250{,}000(0.50)=125{,}000\), and the expected failure count is also 125,000. Both expected counts are at least 10, so the Large Counts condition is met.

Do: Calculate the sample proportion, null standard error, and test statistic:

$$ \hat p=\frac{125{,}750}{250{,}000}=0.503 \qquad SE_0=\sqrt{\frac{0.50(0.50)}{250{,}000}}=0.001 $$
$$ z=\frac{0.503-0.50}{0.001}=3.00 $$

The two-sided p-value is \(2\text{normalcdf}(3.00,1\text{E}99,0,1)\approx0.0027\), rounded. Since \(0.0027\leq0.05\), reject \(H_0\).

Conclude: The data provide convincing evidence that the true renewal proportion differs from 50%; the observed proportion is higher by 0.003, or 0.3 percentage points. If that estimated difference applied to 5,000,000 comparable customers, it would correspond to \(0.003(5{,}000{,}000)=15{,}000\) additional renewals—above the team’s stated 10,000-customer practical benchmark. This makes the estimated difference potentially important in this setting, but the projection is not a guarantee: it assumes the sample represents the larger group and that the observed difference is a good estimate of the population difference. Statistical significance supports evidence of a difference; operational consequences and uncertainty still belong in the decision.

Common Mistakes and AP Exam Tips

A strong interpretation keeps the test decision separate from the judgment about consequences. Name what the test establishes, quantify the observed difference, and explain what context is needed to judge whether it matters.

  • Equating “significant” with “important.” Statistical significance means the p-value is at most \(\alpha\), so the test rejects \(H_0\). It does not mean the difference is large or valuable. State the estimated difference and discuss its context separately.
  • Ignoring sample size. A very large \(n\) can make a tiny difference produce a small p-value. Report the actual difference, not only the test statistic or p-value.
  • Calling a non-significant result proof of no effect. “Fail to reject” means the data do not provide convincing evidence against \(H_0\) at the chosen significance level. It does not establish \(p=p_0\).
  • Using a practical threshold that was never supplied. Do not claim a result is practically important or unimportant based on an invented universal cutoff. Use a threshold stated in the problem or explain what decision-relevant information is missing.
  • Treating a sample estimate as the known population effect. A sample difference can inform a practical discussion, but it is not the exact value of \(p-p_0\). Be cautious when projecting it to a larger group.
  • Misreading percentage points. A change from 50.0% to 50.1% is 0.1 percentage points. Clear units prevent a small difference from sounding larger or smaller than it is.
AP Exam Tip: After the test decision, distinguish the conclusion about evidence from the practical interpretation. For example: “The data provide convincing evidence that the population proportion differs from the benchmark. The observed difference is 0.1 percentage points; whether that matters depends on the consequences and the practical threshold in this setting.”

Key Takeaway

A p-value addresses evidence against a null model, not whether an effect is worth acting on. Large samples can make small differences statistically significant, while smaller samples can leave potentially important differences uncertain. A careful conclusion reports the evidence, the estimated size in context, and the information needed to judge practical importance.

Key takeaway: Statistical significance is determined by comparing the p-value with \(\alpha\), including equality. Practical significance depends on the size and consequences of the difference in context; one does not automatically imply the other.

Check Your Understanding

Use the distinction between statistical evidence and practical importance in each response.

  1. A test has p-value \(0.050\) and uses \(\alpha=0.05\). Is the result statistically significant under the decision rule, and what is the test decision?
  2. A very large sample gives a p-value of \(0.001\), but the observed difference is 0.02 percentage points. What does the p-value establish, and what does it not establish?
  3. A test fails to reject \(H_0:p=0.40\). Explain why it would be incorrect to conclude that \(p\) must equal 0.40.
  4. A sample estimate is 0.7 percentage points above a benchmark, while the organization’s stated practical threshold is 1 percentage point. What can be said about the estimate, and what cannot be concluded about the exact population difference from that fact alone?
  5. If an estimated increase is small for one person but applies to several million people, what contextual quantity could help assess its practical importance?