Tutorials › AP Statistics › Effect Size in Terms of the Original Units

Inference errors and practical significance · Tutorial 782 of 1000

Effect Size in Terms of the Original Units

Express a mean difference in the units people use in the situation, then judge its size against meaningful context and data variability.

Intermediate 9 min read

What You'll Learn

  • Define a raw effect estimate for a mean or mean difference in the original measurement units.
  • Interpret the sign of a difference using the order of the groups or paired measurements.
  • Compare an estimated difference with a context-based benchmark without treating that benchmark as universal.
  • Use a standard deviation as a description of typical spread, not as a measure of practical importance.
  • Read a confidence interval in original units to assess which effect sizes remain plausible.
  • Distinguish the magnitude of an effect from the strength of evidence against a null hypothesis.

Start With the Difference People Can Understand

The previous tutorial, “Is a Statistically Significant Difference in Means Practically Important?”, separated evidence of a difference from the question of whether its size matters. This tutorial focuses on describing that size. A mean difference in the original units—minutes, points, kilograms, or another familiar measurement—can often be understood directly in the setting where it occurs.

For a one-sample problem, the estimated difference from a benchmark is \(\bar{x}-\mu_0\). For two independent groups, it is \(\bar{x}_1-\bar{x}_2\), with the group order stated. For paired data, it is \(\bar{d}\), where the order used to define each difference must be clear. Each estimate has the same units as the original quantitative variable.

Definition: In this tutorial, the raw effect estimate is the estimated mean difference expressed in the original units of the response. Its sign shows the direction of the difference according to the stated order; its magnitude describes how far apart the estimated means are in those units.

For example, an estimated difference of 3 minutes says that the sample means are 3 minutes apart. To judge whether 3 minutes matters, you need context: a change of 3 minutes may be important for one process and negligible for another. There is no universal number of minutes, points, or kilograms that defines an important effect.

A standard deviation adds another useful comparison. It describes the typical spread of individual observations around a mean, or, for paired data, the spread of the individual differences around their mean. Comparing a mean difference with a relevant standard deviation helps describe the difference relative to the variability in that setting. That comparison does not replace the raw-unit description or determine practical importance by itself.

Keep Units, Direction, and Spread Straight

Always state what is being subtracted. If group 1 is subtracted from group 2, reversing the order reverses the sign but not the distance between the sample means. For paired data, define \(d\) in words before interpreting \(\bar{d}\). A positive paired mean could describe an increase or a decrease, depending on that definition.

The standard deviation is in the same units as the response. For a descriptive comparison, you might say that an estimated difference is about one-half of a typical within-group standard deviation. The calculation divides quantities with the same units, so the comparison has no units. But it does not make the original effect easier to act on than a clear statement such as “the estimated difference is 2 points.”

How to compare a difference with spread: Report the raw effect estimate with its units first. Then, if it helps describe the setting, compare its magnitude with a relevant standard deviation. Treat that comparison as descriptive context—not as a universal cutoff for practical importance.

A difference can be small relative to individual variation and still matter for a decision, especially if a small change affects many people or accumulates over time. Conversely, a difference that is large relative to spread is not automatically valuable or important. The consequences and measurement scale in the specific context matter.

Worked Example: A Difference in Setup Time

A fictional workshop compares two setup methods. In a sample of jobs, the mean setup time is 7.6 minutes with Method A and 7.0 minutes with Method B. The sample standard deviations are 1.5 minutes for each method. The workshop’s manager says that reducing average setup time by at least 1 minute would affect the staffing plan. Describe the estimated effect in original units and in relation to spread.

First state the subtraction order: Method A minus Method B. The estimated difference is

$$ \bar{x}_A-\bar{x}_B =7.6-7.0 =0.6\text{ minutes} $$

The sample mean for Method A is 0.6 minutes higher than the sample mean for Method B. If Method B is the newer method intended to reduce time, the estimated reduction is 0.6 minutes—not 1 minute. Relative to the stated practical benchmark, the estimated reduction falls short: \(0.6<1\) minute.

For a simple comparison with the observed spread, divide the magnitude of the difference by either group’s standard deviation. Here both standard deviations are 1.5 minutes:

$$ \frac{|\bar{x}_A-\bar{x}_B|}{s} =\frac{0.6}{1.5} =0.4 $$

So the estimated difference is 0.4 times a within-group standard deviation. This says that the gap between the sample means is smaller than the typical spread of individual setup times within either group. It does not prove that the difference is unimportant; the manager’s staffing benchmark and the consequences of a change are more directly relevant to that judgment.

Interpretation: In this sample, Method B’s mean setup time is 0.6 minutes lower than Method A’s. That estimated reduction is below the manager’s stated 1-minute benchmark and is small relative to the 1.5-minute standard deviation in either group. This is a description of the observed effect, not by itself a conclusion about the population means or the strength of statistical evidence.

A Benchmark and a Standard Deviation Answer Different Questions

A context-based benchmark answers, “How large would a difference need to be to matter for this decision?” A standard deviation answers, “How spread out are individual measurements, or differences, in these data?” These are distinct questions. A benchmark should have a reason grounded in the setting, such as a change in cost, time, performance, or health that would affect a decision.

A difference may be smaller than the standard deviation and still exceed a practical benchmark. It may also be larger than the standard deviation while failing to matter for the decision. When possible, explain both comparisons using words and units rather than relying on a ratio alone.

Worked Example: Weekly Energy Use in Two Buildings

A fictional facilities team records weekly energy use in two buildings with similar operating schedules. Building 1 has a sample mean of 18.4 kilowatt-hours per room per week and a sample standard deviation of 4.0 kilowatt-hours. Building 2 has a sample mean of 16.9 kilowatt-hours per room per week and a sample standard deviation of 3.5 kilowatt-hours. The team considers a reduction of at least 3 kilowatt-hours per room per week large enough to change its energy plan. Describe the sample effect and its scale.

Define the contrast as Building 1 minus Building 2. The estimated difference is

$$ \bar{x}_1-\bar{x}_2 =18.4-16.9 =1.5\text{ kilowatt-hours per room per week} $$

The sample mean is lower in Building 2 by 1.5 kilowatt-hours per room per week. The estimated reduction is below the team’s 3-kilowatt-hour benchmark. For a rough comparison with within-building spread, use the average of the two reported standard deviations:

$$ \frac{4.0+3.5}{2}=3.75 \qquad \frac{1.5}{3.75}=0.40 $$

The mean gap is about 0.40 of this descriptive reference spread. The average standard deviation is being used only as a convenient comparison; it is not a new measure of practical importance and is not a standard error. The raw estimate remains 1.5 kilowatt-hours per room per week.

Interpretation: The sample mean energy use in Building 2 is 1.5 kilowatt-hours per room per week lower than in Building 1. This is below the team’s 3-kilowatt-hour benchmark, and the gap is small relative to the observed within-building standard deviations. These sample summaries describe the size of the observed difference; without an appropriate inference analysis, they do not establish the size of a population mean difference.

Notice that the benchmark gives a decision-focused comparison, while the standard deviations give a variability-focused comparison. Neither comparison tells us by itself how precisely the population difference has been estimated. An estimate based on limited or variable data may have substantial uncertainty, even when its numerical value is clear.

Use an Interval to Describe Uncertainty About the Effect

A confidence interval for a mean or mean difference also stays in the original units. Its center is the sample estimate, and its endpoints show a range of population values that remain plausible under the procedure and its conditions. As covered in “Connecting Confidence Intervals to Test Decisions,” an interval is not a guarantee that the true value lies within it, and it does not turn a practical benchmark into a universal rule.

When judging practical importance, compare the interval—not just its center—with the context benchmark. If all values in the interval are below a threshold, that interval does not support effects as large as the threshold at the stated confidence level. If the interval includes values both below and above the threshold, the estimated effect may be clear in direction while its practical importance remains uncertain.

Worked Example: A Paired Change in Blood Pressure

In a fictional study, 25 adults are randomly selected from a clinic population of 500. Each person’s systolic blood pressure is measured before and after a program. Define \(d=\text{before}-\text{after}\), in millimeters of mercury (mmHg), so a positive difference means a reduction after the program. The sample mean difference is \(\bar{d}=4.8\) mmHg, and the sample standard deviation of the differences is \(s_d=6.0\) mmHg. The differences are roughly symmetric with no apparent outliers. The clinic considers a mean reduction of 5 mmHg important. Estimate the mean effect and use a 95% confidence interval to describe its uncertainty.

State: Let \(\mu_d\) be the true mean before-minus-after systolic blood pressure difference for the population represented by the random sample. Positive values describe an average reduction. The observed raw effect estimate is \(\bar{d}=4.8\) mmHg.

Plan: Use a paired t interval for \(\mu_d\), because each adult provides a linked before-and-after pair. The adults were randomly sampled, and the 10% condition holds because \(25\leq0.10(500)=50\). The differences for distinct adults are treated as independent. The distribution of the differences is roughly symmetric with no apparent outliers, supporting the t procedure.

Do: The standard error is

$$ SE=\frac{s_d}{\sqrt{n}} =\frac{6.0}{\sqrt{25}} =\frac{6.0}{5} =1.2\text{ mmHg} $$

With \(df=25-1=24\), the 95% critical value is approximately \(t^*=2.064\). The interval is

$$ \bar{d}\pm t^*SE =4.8\pm2.064(1.2) =4.8\pm2.4768 \approx(2.32,\ 7.28)\text{ mmHg} $$

The raw estimate is \(4.8\) mmHg, which is \(4.8/6.0=0.80\) of the sample standard deviation of the paired differences. The interval, however, ranges from about 2.32 to 7.28 mmHg. It includes mean reductions below the clinic’s 5-mmHg benchmark and reductions above it.

Conclude: We are 95% confident that the population mean before-minus-after blood pressure difference is between about 2.32 and 7.28 mmHg. The sample estimates a mean reduction of 4.8 mmHg, close to but below the clinic’s 5-mmHg benchmark. Because the interval includes values on both sides of that benchmark, the practical importance of the population mean reduction remains uncertain. The spread comparison describes the effect’s size relative to variability; it does not decide whether the reduction matters clinically.

Common Mistakes and Full-Credit Communication

  • Reporting only a p-value: A p-value describes evidence against a null hypothesis, not the size of the effect. State the estimated difference in the original units as well.
  • Leaving out the subtraction order: “The means differ by 4.8” does not tell the reader which group is higher or what a positive paired difference means. Name the order and interpret the sign in context.
  • Treating a standard deviation as a practical threshold: Spread describes variability, not the smallest change that matters. Use a context-based benchmark for a practical judgment and explain its relevance.
  • Calling a ratio a complete effect description: Saying an effect is 0.4 standard deviations does not tell the reader how many minutes, points, or kilowatt-hours it represents. Give the raw-unit estimate first.
  • Overstating what the sample estimate establishes: A sample difference describes the data. Use an appropriate interval or test, with its conditions checked, to make an inference about population means.
  • Judging an interval only by its center: If an interval includes both values below and values above a practical threshold, say that practical importance remains uncertain rather than treating the point estimate as decisive.

A strong response puts the effect in context: “The estimated mean reduction is 4.8 mmHg. The 95% confidence interval is about 2.32 to 7.28 mmHg, which includes values both below and above the clinic’s 5-mmHg benchmark. Thus the estimated effect is near the benchmark, but its practical importance is not settled by these data.” This identifies the effect, units, uncertainty, and relevant comparison without claiming more than the analysis supports.

Key takeaway: Describe a mean effect first in its original units and with its direction clearly defined. A practical benchmark helps judge whether its size matters in context; a standard deviation describes the effect relative to spread. A confidence interval shows how uncertain the population effect remains.

Check Your Understanding

For each question, identify the effect in its original units and distinguish practical context from spread or uncertainty.

  1. A sample mean is 12.6 minutes with Method A and 11.9 minutes with Method B. Define the difference as A minus B. What is the raw effect estimate, and which method has the lower sample mean?
  2. A two-group comparison estimates a difference of 2.0 points, while a relevant within-group standard deviation is 5.0 points. What is the descriptive ratio of the difference to that standard deviation? What does the ratio not establish?
  3. A team says that a reduction of at least 10 liters per day would affect its water-use plan. Why is that threshold more useful for judging practical importance than a standard deviation alone?
  4. A paired mean change is estimated as 3.2 minutes, with a 95% confidence interval of \((0.4,\ 6.0)\) minutes. A change of 5 minutes is considered important. What does the interval suggest about practical importance?
  5. Why should a conclusion report both the raw-unit effect estimate and, when relevant, a comparison with spread?