Tutorials › AP Statistics › One Group Compared to a Known Value

Choosing a mean-inference procedure · Tutorial 770 of 1000

One Group Compared to a Known Value

Learn to recognize a one-group comparison with a benchmark and state hypotheses about the population mean in the correct direction.

Intermediate 9 min read

What You'll Learn

  • Distinguish one population mean compared with a fixed benchmark from a comparison of two population means.
  • Define \(\mu\) and \(\mu_0\) in context before writing hypotheses.
  • Match a two-sided, greater-than, or less-than alternative to the research question.
  • Check the design, independence, 10% condition, and Nearly Normal condition for a one-sample t test.
  • Interpret a test statistic and p-value for a population mean in context.
  • Avoid using sample statistics or a benchmark value in the wrong hypothesis.

One Sample, One Benchmark

In “A Decision Flowchart for Mean Inference,” you learned to identify the parameter from the research question and the study design. A common mean-inference question uses data from one group and compares its population mean with a fixed benchmark: Is the average fill amount on target? Is the mean device lifetime greater than a stated standard? Is the average waiting time below a goal?

These questions call for a one-sample t test when the response is quantitative and the population standard deviation is not known. The sample provides evidence about one population mean, \(\mu\). The benchmark, written \(\mu_0\), is a particular value specified by a standard, claim, or research question. It is not a second sample or a second population.

Definition: A one-sample t test evaluates a claim about one population mean by comparing the sample mean, \(\bar{x}\), with a fixed benchmark, \(\mu_0\). Its hypotheses are written about the population mean \(\mu\), not the observed sample mean.

The benchmark is treated as fixed for the test. For example, if a machine is supposed to dispense 500 grams per package, then \(\mu_0=500\) grams represents the target. The sample mean might be above or below 500 grams, but that observed statistic does not replace the population parameter in the hypotheses.

Write Hypotheses That Match the Question

The null hypothesis describes the benchmark claim. For a one-sample mean test, it states that the population mean equals the benchmark: \(H_0:\mu=\mu_0\). The alternative hypothesis describes the direction of the research question. It may say the mean differs from the benchmark, is greater than it, or is less than it.

$$ H_0:\mu=\mu_0 \qquad\text{and}\qquad H_a:\mu\ne\mu_0,\ \mu>\mu_0,\ \text{or}\ \mu<\mu_0 $$

Choose the alternative from the question’s wording and purpose, not from the direction of the sample result. “Different from” or “has changed” calls for a two-sided alternative. “Greater than,” “above,” or “increased” calls for \(H_a:\mu>\mu_0\). “Less than,” “below,” or “decreased” calls for \(H_a:\mu<\mu_0\). The equality belongs in the null hypothesis; an alternative hypothesis does not include equality.

Define \(\mu\) in context so there is no doubt about the population and measurement. For instance, “Let \(\mu\) be the true mean fill amount, in grams, for packages produced by this machine during the period of interest.” Then state the benchmark and hypotheses. This identifies what the test can address and prevents a vague hypothesis such as “the average is different.”

Key distinction: A one-sample t test compares one population mean with a fixed value. A two-sample t test compares two population means, while a paired t test uses the population mean of pairwise differences. As in “One-Sample t Versus Two-Sample t,” match the procedure to the target parameter and design.

The test statistic measures how far the sample mean is from the null benchmark in estimated standard-error units. The one-sample t statistic uses the sample standard deviation \(s\), because the population standard deviation is not known.

$$ t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} $$

Here, \(n\) is the sample size. The sign of \(t\) shows whether \(\bar{x}\) is above or below \(\mu_0\). The degrees of freedom for a one-sample t test are \(n-1\). The p-value is calculated in the direction specified by \(H_a\), as in “One-Sided and Two-Sided P-Values Compared.”

Plan Before Calculating

A correct hypothesis pair does not by itself make a t test appropriate. Check the study design and the distributional conditions before relying on the test. For inference about a population mean, an appropriate random sample supports generalizing to the sampled population; randomized assignment in an experiment supports cause-and-effect conclusions but does not by itself make the sample representative. Use the design that justifies the specific inference. If a sample is drawn without replacement from a finite population, check the 10% condition so that observations can reasonably be treated as independent. Independence also depends on how the data were collected; it cannot be established just by inspecting a graph.

For a small sample, inspect the data for strong skewness or outliers. The Nearly Normal condition means the sample data are reasonably compatible with using a t procedure. For a sufficiently large sample, t procedures are generally more robust to skewness, but strong outliers or extreme skewness can still be problematic. These ideas build on “Conditions for a Two-Sample t Test”; for a one-sample test, assess the one sample rather than two separate groups.

Conditions: For generalizing about a population mean, check that the data come from an appropriate random sample; randomized assignment supports causal inference but does not by itself make a sample representative. Check that observations are independent, including the 10% condition when sampling without replacement, and that the sample distribution is reasonably compatible with t inference. For a small sample, look for strong skewness and outliers.

Worked Examples: Matching the Benchmark and Claim

Worked Example: Is Package Fill Different From the Target?

A food producer wants to assess whether a filling machine’s mean package amount differs from its 500-gram target. A random sample of 16 packages has a mean fill of 496.8 grams and a standard deviation of 6.4 grams. Assume the sample’s plot is roughly symmetric with no outliers, and that the production population contains at least 160 packages.

State: Let \(\mu\) be the true mean fill amount, in grams, for packages produced by this machine during the period represented by the sample. The question asks whether the mean differs from 500 grams, so use a two-sided alternative: \(H_0:\mu=500\) grams and \(H_a:\mu\ne500\) grams.

Plan: Use a one-sample t test because there is one quantitative sample and the population mean is being compared with a fixed value. The packages were randomly sampled. The 10% condition is met because \(16/160=0.10\), and the sample is small but has no strong skewness or outliers. The observations can be treated as independent under the stated sampling design, so the conditions support the test.

Do: The estimated standard error is \(s/\sqrt{n}=6.4/\sqrt{16}=1.6\) grams. The test statistic is the sample mean’s distance from the null benchmark divided by this standard error:

$$ t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{496.8-500}{6.4/\sqrt{16}} =\frac{-3.2}{1.6} =-2.00 $$

There are \(16-1=15\) degrees of freedom. For a two-sided test, the p-value is approximately \(0.0639\), rounded. This is the probability, assuming the true mean fill is 500 grams, of obtaining a t statistic at least as far from zero as \(-2.00\) in either direction.

Conclude: At \(\alpha=0.05\), the p-value \(0.0639\) is greater than the significance level, so fail to reject \(H_0\). The sample does not provide convincing evidence that the true mean package fill differs from 500 grams for the production population represented by the random sample.

Worked Example: Is Device Lifetime Greater Than the Benchmark?

A technician investigates whether the mean operating lifetime of a type of rechargeable sensor is greater than 1,200 hours. A random sample of 25 sensors has a mean lifetime of 1,260 hours and a standard deviation of 150 hours. Assume a plot shows no strong skewness or outliers and that the population contains at least 250 sensors.

Identify and state: Let \(\mu\) be the true mean operating lifetime, in hours, for this type of sensor in the population represented by the sample. The claim is that the mean exceeds the benchmark, so \(H_0:\mu=1200\) hours and \(H_a:\mu>1200\) hours. The direction comes from the research question, not from first looking at the sample mean.

Plan: A one-sample t test fits because the response is lifetime, a quantitative variable, and there is one sample compared with a fixed benchmark. The random sample supports the design condition. The 10% condition holds because \(25/250=0.10\). The sample size is 25, and the stated plot gives no strong skewness or outliers; the conditions support t inference.

Do: The estimated standard error is \(150/\sqrt{25}=30\) hours. Thus,

$$ t=\frac{1260-1200}{150/\sqrt{25}} =\frac{60}{30} =2.00 $$

The degrees of freedom are \(25-1=24\). Because the alternative is greater than the benchmark, the p-value is the area to the right of \(t=2.00\), approximately \(0.0285\), rounded. Assuming the population mean lifetime is 1,200 hours, this is the probability of obtaining a t statistic of 2.00 or greater.

Conclude: At \(\alpha=0.05\), \(0.0285<0.05\), so reject \(H_0\). The data provide convincing evidence that the true mean operating lifetime of this type of sensor is greater than 1,200 hours.

Worked Example: Is Delivery Time Below a Service Goal?

A delivery service wants to know whether its mean delivery time is below its 18-minute goal. A random sample of nine deliveries has a mean time of 17 minutes and a standard deviation of 1.5 minutes. Assume the delivery-time plot is reasonably symmetric with no outliers and the population includes at least 90 deliveries.

Define and state: Let \(\mu\) be the true mean delivery time, in minutes, for deliveries in the population represented by the sample. “Below the goal” specifies a less-than alternative. The hypotheses are \(H_0:\mu=18\) minutes and \(H_a:\mu<18\) minutes.

Plan and check: Use a one-sample t test, not a two-sample test: there is one sample and one fixed benchmark. The deliveries were randomly sampled, and \(9/90=0.10\) meets the 10% condition. Because \(n=9\) is small, the shape check matters; the stated plot has no strong skewness or outliers. With independent delivery observations, the conditions support the test.

Do and conclude: The standard error is \(1.5/\sqrt{9}=0.5\) minutes, so

$$ t=\frac{17-18}{1.5/\sqrt{9}} =\frac{-1}{0.5} =-2.00 $$

The degrees of freedom are \(9-1=8\). For the less-than alternative, the left-tail p-value is approximately \(0.0403\), rounded. At \(\alpha=0.05\), reject \(H_0\) because \(0.0403<0.05\). The data provide convincing evidence that the true mean delivery time for the population represented by this sample is below 18 minutes.

Common Mistakes and Full-Credit Communication

  • Writing hypotheses about the sample: \(\bar{x}\) is a statistic that summarizes the sample; it is not the population parameter. Write hypotheses about \(\mu\).
  • Treating the benchmark as a second group: A fixed target such as 500 grams does not describe a second sample. A one-sample test compares \(\mu\) with \(\mu_0\); a two-sample test compares two population means.
  • Putting equality in the alternative: Use equality in \(H_0\). Write \(H_a:\mu\ne\mu_0\), \(H_a:\mu>\mu_0\), or \(H_a:\mu<\mu_0\), according to the question.
  • Choosing the direction after seeing the data: Decide whether the research question is two-sided, greater than, or less than before using the sample result. The observed direction does not justify changing the alternative.
  • Leaving the parameter vague: State what \(\mu\) measures, the population it describes, and the units. “Let \(\mu\) be the average” does not identify the response or population.
  • Skipping conditions: A random sample, an appropriate independence check, the 10% condition when needed, and a shape check are part of the plan. A calculation alone is not a complete test response.

A concise, complete setup might read: “Let \(\mu\) be the true mean lifetime, in hours, for all sensors of this type in the population represented by the random sample. To test whether the mean exceeds 1,200 hours, use \(H_0:\mu=1200\) and \(H_a:\mu>1200\).” This names the parameter, benchmark, direction, and population. The complete response then checks conditions, calculates the test statistic and p-value, and concludes in context, as in “Writing the Full Two-Sample t Test Solution,” with the parameter and procedure adjusted for a one-sample question.

Key takeaway: When one quantitative sample is compared with a fixed benchmark, define the population mean \(\mu\), set \(H_0:\mu=\mu_0\), and choose the alternative direction from the research question. Then check the one-sample t conditions and interpret the evidence about \(\mu\) in context.

Check Your Understanding

For each situation, identify the population mean and benchmark, then write the hypotheses that match the research question.

  1. A random sample of water bottles is used to assess whether the mean volume differs from the labeled 750 milliliters. What are \(H_0\) and \(H_a\)?
  2. A random sample of laptop batteries is used to test whether mean battery life is greater than a 10-hour standard. Define \(\mu\) and write the hypotheses.
  3. A sample of customer support calls is used to investigate whether mean hold time is below a 4-minute goal. What direction should the alternative have?
  4. Why is \(H_a:\bar{x}>1200\) not an appropriate alternative hypothesis for a test about a population mean?
  5. A student sees a sample mean below the benchmark and then changes a planned two-sided alternative to a less-than alternative. What is wrong with that choice?