Tutorials › AP Statistics › Does an Interval for mu d Containing Zero Matter

Paired data and paired t procedures · Tutorial 694 of 1000

Does an Interval for mu d Containing Zero Matter

Use a paired t interval to assess a no-average-difference claim without mistaking failure to reject for proof of no change.

Intermediate 10 min read

What You'll Learn

  • Match a paired t interval’s confidence level to the significance level of a two-sided test.
  • Use zero’s position relative to the interval to decide whether to reject a no-average-difference claim.
  • State what an interval containing zero does and does not establish about the population mean difference.
  • Explain why a one-sided claim or a different significance level may require more information.
  • Handle an interval endpoint exactly at zero using a clearly stated boundary convention.

Why Zero Is the Reference Value

In “Interpreting a Paired t Interval,” you learned to read an interval for the population mean difference, \(\mu_d\), using the subtraction order that defines each paired difference. This tutorial uses that interval to assess a particular claim: that the true mean difference is zero.

For example, if \(d_i=\text{after}_i-\text{before}_i\), then \(\mu_d=0\) means the population mean after-minus-before difference is zero. It is a claim of no average difference under this subtraction order. It does not say that every individual’s two measurements are identical.

Definition: For a paired t interval about \(\mu_d\), zero is the reference value for a no-average-difference claim. The question is whether zero is a plausible value for the population mean difference, given the interval and its confidence level.

For a two-sided test of \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\), a confidence interval and a significance test can give the same decision when they use the same paired differences and match confidence level to significance level. A 95% interval matches a test at \(\alpha=0.05\); a 90% interval matches a test at \(\alpha=0.10\).

$$ \text{Confidence level}=1-\alpha $$

Away from the endpoint boundary, if zero is outside the matching interval, reject \(H_0\). The interval contains only positive or only negative values, so zero is not among the plausible values at that confidence level. If zero is inside the matching interval, fail to reject \(H_0\): zero remains among the plausible values, and the data do not provide convincing evidence of a nonzero population mean difference at the corresponding significance level.

This is a decision rule for a two-sided test. It does not mean an interval containing zero proves that \(\mu_d=0\), or that a treatment had no effect. It means the data and procedure do not rule out zero at the matched confidence and significance levels. The interval may also contain meaningful positive or negative differences.

Connect the Interval to a Two-Sided Test

A paired t interval is calculated from the list of paired differences, using \(\bar d\), \(s_d\), and the number of pairs \(n\). As in “Constructing a Paired t Confidence Interval,” its form is \(\bar d\pm t^*(s_d/\sqrt{n})\), with \(df=n-1\). For the no-difference claim, the comparison value is zero.

$$ \text{Interval for }\mu_d: \quad \bar d\pm t^*\frac{s_d}{\sqrt{n}} $$

The matching two-sided paired t test asks whether the observed sample mean difference is far enough from zero, relative to its standard error, to provide convincing evidence that the population mean difference is not zero. The interval provides a range of plausible values for \(\mu_d\); checking whether zero is in that range gives the corresponding test decision, subject to the boundary convention discussed below.

Decision rule: For a paired t interval with confidence level \(1-\alpha\) and a two-sided test of \(H_0:\mu_d=0\) at significance level \(\alpha\), zero outside the interval corresponds to rejecting \(H_0\); zero inside the interval corresponds to failing to reject \(H_0\). At exact equality with an endpoint, state and apply a consistent boundary convention.

This connection depends on a match: same paired observations, same definition and order of differences, and matching confidence and significance levels. A 95% interval is not the direct interval-based decision rule for a test at \(\alpha=0.10\). Also, looking at whether zero is in a two-sided interval does not by itself answer a one-sided question such as whether the mean difference is greater than zero.

Worked Examples

Worked Example: Battery Runtime After a Software Update

A fictional technician randomly selects 16 devices from a population of 240 devices and measures each device’s battery runtime before and after a software update. Define \(d_i=\text{after}_i-\text{before}_i\), in hours. The differences have no strong skewness or pronounced outliers. A 95% paired t interval is to be constructed from \(\bar d=3.2\) hours and \(s_d=4.0\) hours. Use the interval to assess the two-sided claim that the population mean difference is zero.

State. Let \(\mu_d\) be the true mean after-minus-before difference in battery runtime, in hours, for devices in this population. We test \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\) at \(\alpha=0.05\). The null claim is that there is no average difference in runtime.

Plan. These are paired data because each device is measured twice. The 16 devices were randomly selected, and distinct devices can reasonably be treated as independent. The 10% condition is met because \(16<0.10(240)=24\). There are 16 differences, and the stated shape shows no strong skewness or pronounced outliers, which supports using a paired t procedure. The 95% interval matches a two-sided test at \(\alpha=0.05\).

Do. There are \(n=16\) pairs, so \(df=16-1=15\). For a 95% interval, \(t^*=\text{invT}(0.975,15)\approx2.131\). The standard error and margin of error are

$$ SE=\frac{4.0}{\sqrt{16}}=1.0\text{ hour}, \qquad ME=2.131(1.0)=2.131\text{ hours}. $$

The interval is

$$ 3.2\pm2.131=(1.069,\ 5.331)\text{ hours}, $$

rounded to \((1.07,\ 5.33)\) hours. Zero is outside this matching interval. Therefore, we reject \(H_0\) at the 0.05 significance level.

Conclude. The data provide convincing evidence that the true mean after-minus-before difference in battery runtime is not zero for devices in this population. Because the entire interval is positive, the evidence points toward a higher mean runtime after the update. This inference does not claim that every device’s runtime increased, and random sampling alone does not establish that the update caused the difference.

Worked Example: Discomfort Ratings After a New Stretching Routine

A fictional clinic randomly selects 12 people from 150 people in its target population and records each person’s discomfort rating before and after trying a stretching routine. Define \(d_i=\text{after}_i-\text{before}_i\), in rating points. The differences are roughly symmetric without pronounced outliers. The sample has \(\bar d=-1.1\) points and \(s_d=2.4\) points. Use a 95% paired t interval to assess a two-sided no-average-difference claim.

State. Let \(\mu_d\) be the true mean after-minus-before difference in discomfort rating, in points, for people in the target population. We test \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\) at \(\alpha=0.05\).

Plan. Each person supplies both measurements, so the observations are paired. The people were randomly selected, and different people can reasonably be treated as independent. The 10% condition is met because \(12<0.10(150)=15\). With 12 pairs, we check the differences’ shape; the roughly symmetric distribution without pronounced outliers supports a paired t procedure. The 95% interval matches the 0.05 two-sided test.

Do. Here \(n=12\), so \(df=12-1=11\). For a 95% interval, \(t^*=\text{invT}(0.975,11)\approx2.201\). The standard error and margin of error are

$$ SE=\frac{2.4}{\sqrt{12}}\approx0.6928\text{ points}, \qquad ME=2.201(0.6928)\approx1.5249\text{ points}. $$

The interval is

$$ -1.1\pm1.5249=(-2.6249,\ 0.4249)\text{ points}, $$

rounded to \((-2.62,\ 0.42)\) points. Zero is inside the matching interval, so we fail to reject \(H_0\) at the 0.05 significance level.

Conclude. The data do not provide convincing evidence that the true mean after-minus-before discomfort-rating difference is different from zero for people in the target population. The interval includes a mean decrease, no average difference, and a mean increase. This is not proof that the routine has no effect; it indicates that this study does not establish a nonzero mean difference at this significance level.

Confidence Level and the Endpoint Boundary

Always match the interval to the test you want to assess. If the question is whether zero can be rejected at \(\alpha=0.10\), use a 90% interval. If it is whether zero can be rejected at \(\alpha=0.05\), use a 95% interval. Intervals from the same data can lead to different decisions because a higher-confidence interval is wider.

Worked Example: A 90% Interval Does Not Settle a 5% Test by Itself

For one paired study, suppose the reported 90% interval is \((0.10,\ 3.10)\) points, and the matching 95% interval from the same data is \((-0.24,\ 3.44)\) points. The differences are defined as after minus before. What can be concluded about the two-sided no-average-difference claim at \(\alpha=0.10\) and at \(\alpha=0.05\)?

State. The claim is \(H_0:\mu_d=0\), with the two-sided alternative \(H_a:\mu_d\ne0\). We assess the claim separately at the two stated significance levels.

Plan. The 90% interval matches \(\alpha=0.10\), and the 95% interval matches \(\alpha=0.05\). Use each interval only for its corresponding test level.

Do. Zero is outside the 90% interval, so at \(\alpha=0.10\) we reject \(H_0\). Zero is inside the 95% interval, so at \(\alpha=0.05\) we fail to reject \(H_0\). The two decisions are not contradictory: the evidence meets the less stringent 10% threshold but not the 5% threshold.

Conclude. At the 0.10 level, the data provide convincing evidence of a nonzero mean after-minus-before difference; at the 0.05 level, they do not provide convincing evidence of a nonzero mean difference. The 90% interval alone would not establish the 0.05-level decision; that requires the matching 95% interval or the test result.

There is one technical boundary case. A closed interval includes its endpoints. Thus, if zero is exactly an endpoint, it is literally inside the interval. But under the common test rule to reject when \(p\leq\alpha\), an exact endpoint corresponds to \(p=\alpha\), so the test rejects. To keep the decision consistent, make the boundary exception explicit: if zero is strictly inside, fail to reject; if zero is strictly outside, reject; if zero exactly equals an endpoint, use the test convention and reject under \(p\leq\alpha\).

Worked Example: Zero Exactly at an Endpoint

Suppose a 95% paired t interval, calculated without rounding, has a lower endpoint exactly equal to zero. The interval is for \(\mu_d\), and the corresponding two-sided test uses \(\alpha=0.05\). What decision follows under the rule to reject when \(p\leq\alpha\)?

State. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d\ne0\), with \(\alpha=0.05\).

Plan. The 95% interval matches the 0.05 test. Because zero is exactly on the boundary, use the stated test convention rather than applying the simple inside-versus-outside rule without qualification.

Do. At an exact endpoint, the test statistic is exactly the critical value in absolute value, so the two-sided p-value is exactly \(\alpha=0.05\). Under the stated rule, \(p\leq\alpha\), we reject \(H_0\). Although zero is included as a closed-interval endpoint, the test convention determines the boundary decision.

Conclude. Under the \(p\leq\alpha\) convention, we reject the no-average-difference claim at the 0.05 level in this exact boundary case. In practice, endpoints are usually rounded, so an endpoint displayed as \(0.00\) may not be exactly zero. Use unrounded values or the test p-value before declaring an exact boundary.

Common Mistakes and AP Exam Tips

  • Using the wrong confidence level. A 95% interval directly matches a two-sided test at \(\alpha=0.05\), not one at \(\alpha=0.10\). State the match before making the decision.
  • Treating failure to reject as proof of no difference. Say “fail to reject \(H_0\)” and explain that the data do not provide convincing evidence of a nonzero mean difference. Do not say that the data prove the mean difference is zero.
  • Forgetting the subtraction order. If \(d_i=\text{after}_i-\text{before}_i\), a positive mean difference points to higher after measurements. Name the ordered difference and its units in the conclusion.
  • Making an individual-level claim. The interval concerns \(\mu_d\), the population mean difference. It does not describe every person’s change.
  • Ignoring a rounded endpoint. A displayed zero may be a rounded small positive or negative number. Do not treat it as an exact boundary unless the unrounded value is exactly zero.
  • Applying the rule without the boundary exception. With a closed interval, exact equality puts zero at an endpoint, while the rule \(p\leq\alpha\) rejects when \(p=\alpha\). State the convention and follow it consistently.

For a clear AP response, name the hypotheses and significance level, identify the matching interval, state whether zero is inside, outside, or exactly at an endpoint, and give the test decision in context. If you fail to reject, describe the lack of convincing evidence rather than claiming that no change exists.

Key takeaway: For a two-sided test of \(H_0:\mu_d=0\), use the interval whose confidence level is \(1-\alpha\). Away from the boundary, zero outside the interval means reject and zero inside means fail to reject. An interval containing zero does not prove no average difference; under the \(p\leq\alpha\) convention, an exact endpoint gives \(p=\alpha\) and leads to rejection.

Check Your Understanding

Assume each interval is for the defined population mean difference and comes from a valid paired t procedure.

  1. A 95% interval for after-minus-before differences is \((0.4,\ 2.8)\) hours. What decision does it support for \(H_0:\mu_d=0\) versus \(H_a:\mu_d\ne0\) at \(\alpha=0.05\), and what does the sign suggest?
  2. A 90% interval for a mean paired difference is \((-1.6,\ 0.3)\) points. What is the decision for a two-sided test at \(\alpha=0.10\)? Does the interval prove the population mean difference is zero?
  3. Why must you use a 95% interval, rather than relying on a 90% interval alone, to apply the interval decision rule directly to a two-sided test at \(\alpha=0.05\)?
  4. A 95% interval has an unrounded lower endpoint exactly zero. Under the convention that rejects when \(p\leq\alpha\), what is the corresponding decision at \(\alpha=0.05\)?
  5. With \(d_i=\text{after}_i-\text{before}_i\), an interval entirely below zero supports a mean difference in which direction?