Tutorials › AP Statistics › Statistical Versus Practical Significance in Group Comparisons

Two-proportion hypothesis tests · Tutorial 537 of 1000

Statistical Versus Practical Significance in Group Comparisons

Learn why a small difference can be statistically significant yet practically unimportant, and how a confidence interval helps judge whether a difference is large enough to matter.

Intermediate 9 min read

What You'll Learn

  • Distinguish statistical significance from practical significance in a two-proportion comparison
  • Define a practical-importance threshold in context before interpreting results
  • Use a confidence interval to assess whether plausible differences are large enough to matter
  • Explain what a statistically significant result does and does not establish
  • Recognize when an inconclusive test still leaves room for a practically important difference

Detectable Does Not Always Mean Important

A two-proportion test can provide convincing evidence that two population rates differ. That conclusion does not, by itself, tell us whether the difference matters in a real decision. A change of one percentage point might be important when millions of people are affected, but not worth the cost of a new program in another setting.

As in “Using an Interval and a Test Together,” use the test to assess evidence against a null difference and the interval to see which population differences are plausible. This tutorial adds a practical question: are those differences large enough to matter in the situation? Answering it requires a context-specific standard, not just a p-value.

Definition: Statistical significance means that the data provide enough evidence, at a specified significance level, to reject a null hypothesis. Practical significance asks whether the size of a difference is important enough to matter for a real-world decision.

For two proportions, the difference \(p_1-p_2\) is measured in proportion units. A difference of \(0.05\) is a difference of 5 percentage points. It is not automatically a 5% relative increase: relative change uses a different comparison and depends on the starting proportion.

To assess practical importance, decision makers can identify a practical-importance threshold before looking at the results. For example, they might decide that an increase of at least 5 percentage points would justify the cost of a program. The threshold depends on the setting, costs, benefits, and consequences; there is no universal cutoff for a meaningful difference.

Key distinction: A test of \(H_0:p_1-p_2=0\) evaluates evidence for a difference from zero. It does not test whether the difference is large enough to meet a practical-importance threshold. Compare the estimated difference and its confidence interval with a context-justified threshold to assess practical importance.

A confidence interval is especially useful because it shows uncertainty about the size of the difference, not just whether zero is plausible. If the entire interval is above zero but below a practical threshold, the data support a positive difference that appears too small to meet that threshold. If the interval includes differences on both sides of the threshold, the study may not establish whether the difference is practically important.

Use the Threshold and Interval Together

Suppose a program is considered worth adopting only if it improves a success rate by at least 5 percentage points. Let Group 1 be the program group and Group 2 the comparison group, so positive values of \(p_1-p_2\) favor the program. The threshold is \(0.05\). A test of equality can assess whether the rates differ; a confidence interval can then show whether plausible values are below, above, or on both sides of \(0.05\).

  • If the test rejects equality and the interval is entirely above zero but below the threshold, there is evidence of a difference, but the plausible differences do not reach the chosen practical threshold.
  • If the test fails to reject equality but the interval includes differences larger than the threshold, there is not convincing evidence of a difference, but a practically important difference remains plausible.
  • If the interval is entirely above the threshold, the plausible differences exceed that threshold. This supports practical importance under the stated criterion.
  • If the interval crosses the threshold, the data do not clearly place the difference on one side of the practical criterion.

These are interpretations of the estimate and its uncertainty, not additional hypothesis-test decisions. In particular, do not change the hypotheses after seeing the data to make a result sound practically important. Explain what threshold is being used and why it is relevant in context.

Worked Examples

Worked Example: A Statistically Significant but Practically Small Difference

Question: A fictional company compares two independent random samples of customers who received different registration reminders. In Group 1, 10,400 of 20,000 customers complete registration; in Group 2, 10,000 of 20,000 do. The company decided in advance that an increase of at least 5 percentage points would be needed to justify the more expensive reminder. At \(\alpha=0.05\), test for a difference and construct a 95% confidence interval. Then assess practical importance.

State: Let \(p_1\) be the true proportion of customers receiving Reminder 1 who complete registration, and \(p_2\) the corresponding proportion for Reminder 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The practical-importance threshold for a positive difference is \(0.05\).

Plan: The groups are independent random samples, and both use the same registration outcome. Assume each customer population has at least 200,000 customers, so each sample of 20,000 is at most 10% of its population; the 10% condition is met. For the test, the pooled proportion is \((10{,}400+10{,}000)/40{,}000=0.51\). In each group, expected successes under the null are \(20{,}000(0.51)=10{,}200\), and expected failures are \(20{,}000(0.49)=9{,}800\); all four expected counts are at least 10, so the Large Counts condition is met. For the interval, the observed successes and failures are 10,400 and 9,600 in Group 1 and 10,000 and 10,000 in Group 2; all are at least 10. The conditions support both procedures.

Do: The sample proportions are \(\hat{p}_1=10{,}400/20{,}000=0.52\) and \(\hat{p}_2=10{,}000/20{,}000=0.50\), so the estimated difference is \(0.02\), or 2 percentage points. The pooled standard error and test statistic are:

$$ SE_{\text{pooled}} =\sqrt{0.51(0.49)\left(\frac{1}{20{,}000}+\frac{1}{20{,}000}\right)} =\sqrt{0.00002499} \approx0.004999 $$
$$ z=\frac{0.52-0.50}{0.004999}\approx4.001 $$

The two-sided p-value is approximately \(0.0001\) (more precisely, about \(0.000063\)). Since this is less than \(0.05\), reject \(H_0\). The data provide convincing evidence that the true registration proportions differ.

For the 95% confidence interval, use the unpooled standard error:

$$ SE_{\hat{p}_1-\hat{p}_2} =\sqrt{\frac{0.52(0.48)}{20{,}000}+\frac{0.50(0.50)}{20{,}000}} =\sqrt{0.00002498} \approx0.004998 $$

Using \(z^*=1.96\), the interval is \(0.02\pm1.96(0.004998)\), or approximately \((0.0102,\ 0.0298)\). In context, plausible values for the difference in registration proportions run from about 1.02 to 2.98 percentage points, with Group 1 higher.

Conclude: The difference is statistically significant at the 0.05 level, but the entire confidence interval is below the company’s 5-percentage-point practical threshold. Under that stated criterion, the data indicate a real difference that is too small to justify the more expensive reminder. Statistical significance does not make a difference practically important.

Worked Example: A Nonsignificant Result That Does Not Rule Out Practical Importance

Question: In a fictional randomized experiment, 200 patients are randomly selected from a pool of more than 2,000 eligible patients and assigned equally to two appointment-reminder methods. In Group 1, 24 of 100 patients attend; in Group 2, 16 of 100 attend. The clinic considers an increase of 10 percentage points practically important. At \(\alpha=0.05\), test for a difference, construct a 95% confidence interval, and assess what the results say about that threshold.

State: Let \(p_1\) and \(p_2\) be the true proportions of eligible patients who would attend under Reminder 1 and Reminder 2, respectively. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The estimated difference is compared with the practical threshold of \(0.10\).

Plan: Patients were randomly selected and randomly assigned, and each patient received only one reminder, so the groups are independent. The sample of 200 is less than 10% of the eligible pool of more than 2,000, satisfying the 10% condition for sampling. The pooled proportion is \((24+16)/200=0.20\). Under the null, each group has 20 expected successes and 80 expected failures, so all four expected counts are at least 10. For the interval, the observed successes and failures are 24 and 76 in Group 1 and 16 and 84 in Group 2; all are at least 10. Thus the conditions for the test and interval are met.

Do: The sample proportions are \(0.24\) and \(0.16\), giving an estimated difference of \(0.08\), or 8 percentage points. The pooled standard error is:

$$ SE_{\text{pooled}} =\sqrt{0.20(0.80)\left(\frac{1}{100}+\frac{1}{100}\right)} =\sqrt{0.0032} \approx0.05657 $$

The test statistic is \(z=0.08/0.05657\approx1.414\), giving a two-sided p-value of approximately \(0.1573\). Because \(0.1573>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the attendance proportions differ.

For the 95% confidence interval, the unpooled standard error is:

$$ SE_{\hat{p}_1-\hat{p}_2} =\sqrt{\frac{0.24(0.76)}{100}+\frac{0.16(0.84)}{100}} =\sqrt{0.003168} \approx0.05628 $$

The interval is \(0.08\pm1.96(0.05628)\), or approximately \((-0.0303,\ 0.1903)\). It includes zero and also includes differences above the clinic’s 10-percentage-point threshold.

Conclude: The test does not establish a difference, but the interval is wide enough to include a practically important increase for Reminder 1, as well as a small decrease. The results do not establish that the reminder is practically helpful or that it is practically unimportant. More precise data would be needed to distinguish those possibilities. Failing to reject equality is not evidence that the two methods have identical effects.

Worked Example: Evidence That a Difference Exceeds the Practical Threshold

Question: In a fictional randomized training study, 500 employees are assigned to each of two formats. A defined skills check is passed by 180 employees in Group 1 and 125 in Group 2. The organization decided that an improvement of at least 5 percentage points would be meaningful. At \(\alpha=0.05\), test whether the rates differ and use a 95% confidence interval to assess practical importance.

State: Let \(p_1\) be the true proportion of employees who would pass under Format 1, and \(p_2\) the corresponding proportion under Format 2. Test \(H_0:p_1=p_2\) against \(H_a:p_1\ne p_2\). The practical-importance threshold is \(0.05\).

Plan: Employees were randomly assigned to independent groups, and all were assessed using the same pass criterion. Assume the 1,000 employees were randomly selected from more than 10,000 eligible employees, so the sample is at most 10% of that population. The pooled proportion is \((180+125)/1{,}000=0.305\). Under the null, each group has 152.5 expected successes and 347.5 expected failures, all at least 10. For the interval, the observed successes and failures are 180 and 320 in Group 1 and 125 and 375 in Group 2, all at least 10. The conditions for both procedures are met.

Do: The sample proportions are \(180/500=0.36\) and \(125/500=0.25\), so the estimated difference is \(0.11\), or 11 percentage points. The pooled standard error and test statistic are:

$$ SE_{\text{pooled}} =\sqrt{0.305(0.695)\left(\frac{1}{500}+\frac{1}{500}\right)} =\sqrt{0.0008479} \approx0.02912 $$
$$ z=\frac{0.11}{0.02912}\approx3.778 $$

The two-sided p-value is approximately \(0.0002\) (about \(0.00016\)). Since it is less than \(0.05\), reject \(H_0\). There is convincing evidence that the true pass proportions differ.

For the 95% interval, the unpooled standard error is:

$$ SE_{\hat{p}_1-\hat{p}_2} =\sqrt{\frac{0.36(0.64)}{500}+\frac{0.25(0.75)}{500}} =\sqrt{0.0008358} \approx0.02891 $$

The interval is \(0.11\pm1.96(0.02891)\), or approximately \((0.0533,\ 0.1667)\). The entire interval is above the 5-percentage-point threshold.

Conclude: The test provides convincing evidence that the pass proportions differ, and the interval indicates that Format 1’s proportion is higher. Because even the lower endpoint is above the organization’s practical threshold, the data support a difference that is practically important under the organization’s stated criterion. This conclusion depends on that criterion being appropriate for the decision.

Common Mistakes and AP Exam Tips

  • Equating a small p-value with a large effect: A p-value measures how unusual the data are under the null model; it does not measure the size or importance of the difference. Report the estimated difference and, when available, the interval.
  • Calling a difference important without a reason: A threshold should come from the context, such as costs, benefits, or a decision rule. “Five percentage points is important” is not self-justifying; explain why that standard matters in the situation.
  • Treating a nonsignificant result as proof of no meaningful difference: Check the interval. If it includes effects larger than the practical threshold, those effects remain plausible even when the test fails to reject.
  • Mixing up percentage points and percent change: A difference between proportions of \(0.30\) and \(0.25\) is 5 percentage points. Do not describe it as a 5% increase without calculating and defining relative change.
  • Changing the threshold after seeing the result: Choosing a standard to make the observed result appear important weakens the argument. State the practical criterion and its context clearly, ideally before interpreting the data.
  • Claiming the same scope for every study: As explained in “Scope of Inference Based on How Data Were Collected,” use random sampling and random assignment to determine whether generalization or cause-and-effect conclusions are justified.
AP Exam Tip: Keep the statistical conclusion and practical judgment separate. First compare the p-value with \(\alpha\) and state whether you reject or fail to reject \(H_0\). Then report the estimated difference and interpret the confidence interval relative to a justified practical threshold. Conclude only what the design and results support.

Key Takeaway

Statistical significance is about evidence against a null difference; practical significance is about whether the size of a difference matters in context. A confidence interval helps connect the two by showing plausible effect sizes. A statistically significant difference can be too small to matter, and a nonsignificant result can still leave practically important differences plausible.

Key takeaway: Define a context-appropriate practical threshold, then compare it with the estimated difference and confidence interval. Do not use the p-value alone to judge importance, and do not mistake failure to reject for proof that no meaningful difference exists.

Check Your Understanding

For each situation, distinguish evidence of a difference from evidence that the difference matters in practice.

  1. A 95% confidence interval for \(p_1-p_2\) is \((0.012,\ 0.028)\), and the stated practical threshold is \(0.05\). What does the interval suggest about the direction and practical size of the difference?
  2. A test at \(\alpha=0.05\) has p-value \(0.18\), while the 95% interval includes values above a stated practical threshold. What can and cannot be concluded?
  3. Why does a small p-value not, by itself, show that a two-proportion difference is important in practice?
  4. A difference is 0.04 between two proportions. Express it in percentage points, and explain why that is not automatically a 4% relative increase.
  5. Why should a practical-importance threshold be justified by the context rather than chosen only after looking at the sample results?