Two Different Questions About a Difference in Means
When a t test finds a statistically significant difference in means, it gives evidence against the null hypothesis. That does not automatically mean the difference matters in practice. A difference can be statistically convincing but too small to affect a decision; another difference can be large enough to matter but estimated too imprecisely to be statistically significant.
This distinction applies to one-sample, paired, and two-sample t procedures. The relevant population parameter might be a population mean \(\mu\), a mean of paired differences \(\mu_d\), or a difference between two population means \(\mu_1-\mu_2\). In every case, the test evaluates evidence about a population parameter. Practical importance asks whether the size of the difference would matter in the particular setting.
As explained in “What a P-Value Says About Sample Means,” a p-value describes how unusual the test statistic would be if the null hypothesis were true. It is not a measure of how large or useful the observed difference is. The p-value also depends on the standard error, which is affected by sample size and variability. A small difference can therefore produce a small p-value when the estimate is sufficiently precise.
A Small Difference, a Small P-Value
Suppose a production manager considers a change of at least 0.50 minutes in average processing time large enough to affect staffing plans. That 0.50-minute threshold is a contextual benchmark, not a universal statistical rule. The manager should set such a benchmark using the consequences of the decision, ideally before examining the sample results.
Worked Example: Statistically Significant but Not Practically Important
A fictional plant randomly samples 100 machine cycles from 2,400 cycles completed during a production period. Processing time is measured in minutes. The sample mean is \(\bar{x}=10.20\) minutes, and the sample standard deviation is \(s=0.590\) minutes. The manager asks whether the population mean processing time differs from the 10.00-minute target. A change of 0.50 minutes or more in either direction would be considered practically important for staffing. The sample distribution is roughly symmetric with no apparent outliers. Use \(\alpha=0.05\).
State: Let \(\mu\) be the true mean processing time, in minutes, for cycles in this production period. The hypotheses are \(H_0:\mu=10.00\) and \(H_a:\mu\ne10.00\). The sample difference from the target is \(10.20-10.00=0.20\) minutes.
Plan: Use a one-sample t test for \(\mu\). The response is quantitative, and one random sample is being compared with a fixed target. The cycles were randomly sampled, and the 10% condition is satisfied because \(100\leq0.10(2400)=240\). The roughly symmetric distribution with no apparent outliers supports a one-sample t procedure. The sample observations are treated as independent.
Do: The standard error and test statistic are
The degrees of freedom are \(100-1=99\). The two-sided p-value is approximately \(0.0010\), rounded. Since \(0.0010<0.05\), reject \(H_0\). The 95% confidence interval uses \(t^*\approx1.984\) with 99 degrees of freedom:
Conclude: The data provide convincing evidence that the population mean processing time differs from 10.00 minutes. We are 95% confident that the mean is between about 10.083 and 10.317 minutes. Although the difference is statistically significant, the entire interval indicates a change smaller than the manager’s 0.50-minute practical-importance threshold. The result is evidence of a difference, but not of a difference large enough to meet that stated threshold.
This conclusion uses the threshold as a practical guide; it does not prove that the change has no consequences. A different workplace might have a different threshold because its staffing costs or production demands differ. Also, statistical significance here is not a statement about each individual cycle: the parameter is the population mean.
The example shows why sample size matters to significance. The estimate is only 0.20 minutes from the target, but \(n=100\) and the observed variability make the standard error small. The test statistic is the estimated difference measured in standard errors, so a modest difference can be far enough from zero to produce a small p-value. Neither the p-value nor the word “significant” replaces examining the estimated difference in its original units.
Use the Confidence Interval to See What Sizes Fit the Data
A confidence interval gives a range of plausible values for the population mean or mean difference. Its location and width help separate statistical evidence from practical judgment. If a two-sided test rejects a null value at level \(\alpha\), the matching \(100(1-\alpha)\%\) confidence interval excludes that null value, as covered in “Connecting Confidence Intervals to Test Decisions.” But the interval may still contain values that are too small to matter in practice.
For a practical-importance threshold, examine whether the interval lies below, above, or across that threshold. If the interval is entirely inside a range of differences considered negligible, the data support the view that the plausible differences are small relative to that benchmark. If the interval includes both negligible and important values, the study has not pinned down practical importance clearly. An interval is not a formal proof that every value inside it is equally likely, nor does it establish exact equality.
A Possibly Important Difference Can Be Statistically Unclear
The reverse situation is also possible: the sample difference may be large enough to matter, but the evidence may not be strong enough to distinguish it from no difference. Small samples or substantial variability can produce a wide interval. In that situation, “not statistically significant” does not mean “not practically important.” It means the data do not provide convincing evidence against the null at the selected significance level.
Worked Example: A Potentially Important Difference With Uncertainty
In a fictional randomized experiment, 16 volunteers are assigned equally to one of two study-room lighting settings. Each person contributes one independent score, measured as hours of focused study during a week. Setting 1 has \(n_1=8\), \(\bar{x}_1=34\) hours, and \(s_1=5\) hours. Setting 2 has \(n_2=8\), \(\bar{x}_2=31\) hours, and \(s_2=5\) hours. The score distributions are roughly symmetric with no apparent outliers. The team would consider a difference of 4 hours per week practically important. Test whether the population means differ using \(\alpha=0.05\).
State: Let \(\mu_1\) and \(\mu_2\) be the true mean focused-study hours for the populations represented by volunteers using settings 1 and 2. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Plan: Use an unpooled two-sample t test for \(\mu_1-\mu_2\). The response is quantitative, and two independent groups of different volunteers are being compared. Participants were randomly assigned, each contributed one score, and the group distributions are roughly symmetric with no apparent outliers. There is no 10% condition to check because the description gives a randomized experiment, not sampling without replacement from a stated finite population.
Do: The observed difference is \(34-31=3\) hours. The unpooled standard error and test statistic are
The unpooled degrees of freedom are \(14\). The two-sided p-value is approximately \(0.2501\), rounded. Since \(0.2501>0.05\), fail to reject \(H_0\). A matching 95% confidence interval uses \(t^*\approx2.145\):
Conclude: The data do not provide convincing evidence that the population mean focused-study hours differ between the two lighting settings. The estimated difference is 3 hours, which is below the team’s 4-hour practical threshold, but the interval includes differences larger than 4 hours as well as differences in the opposite direction. The results leave practical importance uncertain; they do not show that the settings are equivalent or that an important difference is impossible.
Because volunteers were randomly assigned, the experiment supports a cause-and-effect comparison for the participants in the experiment. Random assignment alone does not justify generalizing the result to a broader population of people. The interval and conclusion should not be presented as applying to all students unless the sampling design supports that generalization.
When the Difference Is Both Significant and Important
A result can also provide evidence of a difference that is large enough to matter. For paired data, define the order of the difference before interpreting its sign. The t procedure then estimates or tests the mean of those pairwise differences, not the two measurements as though they came from unrelated groups.
Worked Example: A Mean Change That Exceeds a Practical Threshold
A fictional usability team randomly selects 16 technicians from a workforce of 200. Each technician completes a timed task using two dashboard layouts, with the order randomized. Let \(d=\text{time with the old layout}-\text{time with the new layout}\), in minutes, so positive differences indicate less time with the new layout. The sample mean time is 18.6 minutes with the old layout and 16.2 minutes with the new layout, giving \(\bar{d}=2.4\) minutes. The sample standard deviation of the differences is \(s_d=2.4\) minutes. The differences are roughly symmetric with no apparent outliers. The team considers a mean reduction of at least 1 minute practically important.
Identify and check: The same technicians used both layouts, so this is paired data. The target is \(\mu_d\), the true mean time with the old layout minus the true mean time with the new layout for the population represented by the random sample. Use a paired t procedure. The technicians were randomly sampled, the 10% condition holds because \(16\leq0.10(200)=20\), and the differences are roughly symmetric with no apparent outliers. Randomizing the layout order helps address order effects.
Calculate: The standard error is \(s_d/\sqrt{n}=2.4/\sqrt{16}=0.6\) minutes. The test of \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\) gives
The two-sided p-value is approximately \(0.001159\), rounded, so the result is statistically significant at \(\alpha=0.05\). For practical interpretation, a 95% confidence interval uses \(t^*\approx2.131\):
Interpret: We are 95% confident that the population mean time with the old layout minus the population mean time with the new layout is between about 1.12 and 3.68 minutes. The interval is entirely above the 1-minute practical threshold. The data therefore provide evidence of a statistically significant difference, and the plausible mean reductions all exceed the team’s stated practical-importance benchmark. Because the technicians were randomly sampled and the layout order was randomized, the study supports generalizing to the workforce represented by that sample and attributing the comparison to the layouts, subject to the study’s design and conditions.
Common Mistakes and Full-Credit Communication
- Calling a small p-value a large effect: A p-value describes evidence against \(H_0\), not the size or usefulness of the difference. Report the estimated difference in its original units.
- Assuming “not significant” means “no difference”: When \(p>\alpha\), fail to reject \(H_0\). Do not claim the means are equal or that a meaningful difference has been ruled out. Use the interval to describe the range of plausible values.
- Using a practical threshold without context: A cutoff such as 0.50 minutes or 4 hours is meaningful only if it has a reason in that setting. Explain what decision or consequence makes the threshold relevant.
- Ignoring the direction and order of a difference: For paired data, define \(d\); for two groups, state which group is group 1. Then interpret the sign and interval endpoints using that order.
- Claiming that statistical significance proves practical importance: A significant result may be far smaller than the amount that matters. Compare the confidence interval with the practical threshold before making that claim.
- Overstating what an interval establishes: An interval entirely below a threshold supports the conclusion that the plausible values are below it at that confidence level. It does not prove a population difference is exactly zero or that no consequence exists.
- Confusing random assignment with random sampling: Random assignment supports cause-and-effect conclusions about the experimental units; random sampling supports generalization to a population. State only the conclusion the design warrants.
A complete response separates the statistical decision from the practical judgment. For example: “The test provides convincing evidence that the population mean differs from the target. The confidence interval estimates the difference as small relative to the stated 0.50-minute threshold, so statistical significance does not imply practical importance in this setting.” If the interval is wide, say that practical importance remains uncertain rather than forcing a yes-or-no claim.
Check Your Understanding
For each situation, separate the statistical evidence from the practical judgment. Use the stated context and units in your reasoning.
- A one-sample t test gives \(p=0.004\) for a difference from a target. The estimated difference is 0.15 kilograms, and the 95% confidence interval is \((0.05,\ 0.25)\) kilograms. If changes of at least 0.8 kilograms matter operationally, what does the result say about statistical significance and practical importance?
- A paired t test has \(p=0.12\). Its 95% confidence interval for mean change is \((-0.4,\ 2.8)\) minutes, and a change of 2 minutes is considered important. Can you conclude that the process has no practically important effect? Explain.
- Two independent groups have sample means of 72 and 68 points. The test is statistically significant, and the 95% interval for the mean difference is \((1,\ 7)\) points. What additional contextual information is needed to judge practical importance?
- Why can a very large sample make a small mean difference statistically significant? Name the quantity in the t statistic that helps explain this.
- A randomized experiment uses volunteers but not a random sample from a stated population. What kind of conclusion can random assignment support, and what kind of conclusion does it not establish by itself?