Tutorials › AP Statistics › Common Errors When Selecting a Procedure

Choosing a mean-inference procedure · Tutorial 778 of 1000

Common Errors When Selecting a Procedure

Use a design-first checklist to spot procedure-selection errors and justify the mean-inference method that fits the data and research question.

Intermediate 9 min read

What You'll Learn

  • Distinguish a single population mean from a mean of paired differences and a difference between two independent population means.
  • Use the link between observations—not the number of columns or equal sample sizes—to identify paired data.
  • Choose a confidence interval or test based on whether the question asks for an estimate or evidence about a claim.
  • Check that the procedure’s conditions apply to the actual observations being analyzed.
  • Explain why a tempting alternative procedure does not match the study design.

A Short Audit Before You Choose

In “Why the Wrong Procedure Gives Misleading Results,” you saw how ignoring genuine pairs can change a standard error and the apparent evidence. A useful next step is to make procedure selection an explicit audit rather than a quick guess from a table or calculator menu. The key is to connect the question, the population parameter, and the way the observations were obtained.

For mean inference, the main choices are a one-sample t procedure, a paired t procedure, and an unpooled two-sample t procedure. After identifying which one fits, decide whether the goal is an interval estimate or a test of a claim. These are separate decisions: the design determines the kind of t procedure, while the wording of the question determines interval versus test.

Key idea: First identify the quantitative response and the population quantity being studied. Then trace how each observation is linked—or not linked—to other observations. Finally, decide whether the question asks to estimate that quantity or to assess evidence about a claim.

This audit builds on “Identifying the Parameter in a Mean Problem,” “One-Sample t Versus Two-Sample t,” “Paired t Versus Two-Sample t,” and “Interval or Test: What Is the Question Asking.” Here, the emphasis is on the clues that prevent common misclassifications and on explaining why a plausible-looking alternative is wrong.

Use the Design-First Audit

1
Confirm a quantitative response.
Mean procedures analyze a numerical measurement or count for each observational unit. If the response is categorical, a mean procedure is not the right inference family.
2
Name the target parameter.
Ask whether the question concerns one population mean, a population mean difference for linked observations, or the difference between two population means for independent groups.
3
Trace the study design.
Look for a fixed comparison value, repeated measurements on the same units, deliberate matching, or separate unlinked groups. The design—not the number of columns—determines whether observations are paired.
4
Read the goal and check conditions.
“Estimate” or “how much, on average” points to a confidence interval. “Evidence,” “different,” “greater,” or “less” points to a test. Then check the conditions for the procedure selected.

The population parameter is not the sample statistic. For example, a sample mean such as \(\bar{x}\) describes the observed sample; \(\mu\) is the population mean that inference addresses. For paired data, the parameter is \(\mu_d\), where the difference \(d\) must be defined in a stated order. For two independent groups, the target is a difference such as \(\mu_1-\mu_2\), with the group order made clear.

A common trap is to count groups before identifying the link between observations. “Two sets of numbers” could be repeated measurements on the same people, deliberately matched subjects, or two unrelated samples. Only the first two situations contain genuine pairs. As discussed in “Spotting Paired Designs in Word Problems,” a pair is created by the study design, not by putting values on the same row after data collection.

Worked Example: One Sample Compared with a Fixed Benchmark

A fictional equipment team randomly selects 16 rechargeable lamps from a shipment of 500. The team measures each lamp’s runtime in hours. The sample mean is 12.8 hours, and the sample standard deviation is 1.6 hours. The distribution of the 16 runtimes is described as roughly unimodal with no apparent outliers. The question is whether the population mean runtime exceeds the 12-hour benchmark. Use \(\alpha=0.05\).

State: Let \(\mu\) be the true mean runtime, in hours, for the population represented by the sampled lamps. The claim concerns whether that one population mean is greater than a fixed value. The hypotheses are \(H_0:\mu=12\) and \(H_a:\mu>12\).

Plan: Use a one-sample t test. There is one quantitative sample, and the 12-hour value is a fixed benchmark—not a second sample. The lamps were randomly selected. Since \(16\leq0.10(500)=50\), the 10% condition is satisfied, supporting independence when sampling without replacement. The stated unimodal shape with no apparent outliers supports using a t procedure for this sample of 16.

Do: The standard error is \(s/\sqrt{n}\), so the test statistic is

$$ SE=\frac{1.6}{\sqrt{16}}=0.4 \qquad t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{12.8-12}{0.4}=2.00 $$

The degrees of freedom are \(16-1=15\). The right-tailed p-value for \(t=2.00\) with 15 degrees of freedom is approximately \(0.0320\), rounded. Since \(0.0320<0.05\), reject \(H_0\).

Conclude: The data provide convincing evidence that the population mean runtime for the lamps represented by this sample exceeds 12 hours.

Why the alternatives are wrong: A two-sample t procedure would require a second independent sample, which is not present. A paired t procedure would require two linked measurements per lamp, which are not present either. The benchmark is a constant, not a set of observed runtimes.

Clues That Separate Paired and Independent Groups

Equal sample sizes do not establish pairing. Two independent groups might happen to have the same number of observations; that does not create a unit-to-unit link. Conversely, a paired data set may be stored in two columns, but the analysis uses one difference for each pair. The design description tells you which interpretation is justified.

Another error is to treat “before and after” as a special formula rather than evidence about the design. When each unit contributes a before measurement and an after measurement, the observations are linked. When two unrelated groups are measured once each, they are not. “Two-Sample Versus Paired Test on the Same Numbers” shows why the same-looking numerical summaries can call for different procedures when the design differs.

Worked Example: An Interval for Mean Savings from Paired Measurements

A fictional energy team randomly selects 8 machines from a population of 100. Each selected machine is tested using both an old and a new setting. Define \(d=\text{old-setting energy use}-\text{new-setting energy use}\), so positive differences represent savings with the new setting. The observed differences, in energy units, are \(1,2,2,3,3,4,4,5\). The team asks for an estimate of the population’s mean savings.

Identify the procedure: Each machine supplies both measurements, so the observations are paired. The target is \(\mu_d\), the true mean savings per machine, not two unrelated population means. Because the question asks for an estimate rather than evidence about a claim, use a paired t confidence interval.

Check conditions: The machines were randomly selected, and \(8\leq0.10(100)=10\), so the 10% condition is met. Differences from distinct machines are treated as independent. The eight differences are symmetric around 3 and show no apparent outliers, supporting a t interval for \(\mu_d\).

Calculate: The differences sum to 24, giving \(\bar{x}_d=24/8=3\). Their deviations from 3 are \(-2,-1,-1,0,0,1,1,2\); the squared deviations sum to 12. Thus,

$$ s_d=\sqrt{\frac{12}{8-1}}=\sqrt{\frac{12}{7}}\approx1.309 \qquad SE=\frac{s_d}{\sqrt{8}}\approx0.463 $$

For a 95% interval with \(7\) degrees of freedom, \(t^*\approx2.365\). The margin of error is \(2.365(0.463)\approx1.095\), so the interval is

$$ \bar{x}_d\pm t^*SE =3\pm1.095 \approx(1.905,\ 4.095) $$

Interpret: We are 95% confident that the population mean energy savings with the new setting is between about 1.91 and 4.09 energy units per machine. The interval estimates a mean difference; it does not say that every machine saves that amount.

Why the alternatives are wrong: A two-sample t interval would discard the machine-by-machine links and treat the settings as separate groups. A one-sample t interval on the new-setting measurements alone would not estimate savings relative to the old setting. Also, a test is not the requested output: the question asks for an estimate, not evidence for a directional or nonzero claim.

When Two Independent Groups Really Are Present

Use an unpooled two-sample t procedure when the question compares the means of two independent groups and there is no genuine one-to-one pairing. “Independent” here describes the design connection between the samples; it does not mean that the two sample sizes must differ or that the groups must be equal in every other feature. In a randomized experiment, random assignment can create independent treatment groups. In a sampling study, the groups may be separately selected.

For this procedure, check the quantitative response, the random process, independence within and between groups, the 10% condition when sampling without replacement, and the shape of each group’s data. As “Matching Procedures to Conditions” emphasizes, inspect the observations relevant to the chosen method: for a two-sample t procedure, consider each group separately.

Worked Example: Two Unlinked Treatment Groups

In a fictional experiment, 20 volunteers are randomly assigned to one of two independent training plans, with 10 volunteers in each group. After four weeks, each volunteer completes a quantitative endurance task measured in points. Plan A has a sample mean of 72 points and a sample standard deviation of 8 points. Plan B has a sample mean of 66 points and a sample standard deviation of 7 points. Each volunteer follows only one plan, and the researchers ask whether the population mean scores differ.

Identify the procedure: The response is quantitative, and the target is \(\mu_A-\mu_B\), the difference in the two population mean scores. The volunteers are distinct across groups, with no repeated measurements or deliberate matching. The question asks whether the means differ, so the appropriate method is an unpooled two-sample t test with \(H_0:\mu_A-\mu_B=0\) and \(H_a:\mu_A-\mu_B\ne0\).

Check conditions: Volunteers were randomly assigned to the plans, supporting a comparison of treatment groups. Each volunteer contributes one response to one group, so the groups are independent by design. The groups consist of different people, so there is no pairing to preserve. Since this is random assignment rather than sampling from a stated finite population, a 10% sampling condition is not the relevant justification here. For the t procedure, the score distributions in both groups should be reasonably well behaved; assume plots show no strong skewness or apparent outliers.

Show what the selected procedure uses: The unpooled standard error is calculated from each group’s sample standard deviation and sample size:

$$ SE=\sqrt{\frac{s_A^2}{n_A}+\frac{s_B^2}{n_B}} =\sqrt{\frac{8^2}{10}+\frac{7^2}{10}} =\sqrt{11.3} \approx3.362 $$

The observed difference, in the stated order, is \(72-66=6\) points. The sample summaries therefore enter a two-sample t test through the difference between sample means and the unpooled standard error. There is no paired difference for each volunteer because no volunteer received both plans and no volunteers were matched into pairs.

Why tempting alternatives are wrong: Equal group sizes do not make this a paired design. A one-sample t test would compare one group’s mean with a fixed value, but the question compares two group means. A paired t test would require actual linked pairs and would analyze their individual differences instead.

Common Misclassifications and Full-Credit Explanations

  • Calling a benchmark a second sample: If one sample is compared with a stated target such as 12 hours, that target is fixed. It does not provide observations for a two-sample procedure. Name the one population mean and use one-sample t.
  • Calling equal-sized groups paired: Matching sample sizes are not evidence of matching units. State what created the link—same unit twice or deliberate matching—or explain that the groups consist of distinct, unlinked units.
  • Choosing from the table layout: Two columns can hold paired measurements, but they can also list unrelated groups. Identify the observational unit and trace where each value came from before selecting a procedure.
  • Choosing a test when the task asks for an estimate: The design may identify the correct t family, but the request still determines interval versus test. An estimate of a mean or mean difference calls for a confidence interval; a claim and evidence question calls for a test.
  • Skipping the parameter: Naming “paired t” without defining the difference leaves its meaning unclear. Define \(d\) in order and identify \(\mu_d\); for two groups, state the order of \(\mu_1-\mu_2\).
  • Assuming the calculator validates the choice: Software can calculate a result for a selected procedure, but it cannot determine whether the data were paired, whether the response is quantitative, or whether the conditions fit. Make those decisions first.
  • Confusing procedure choice with causal conclusions: Random assignment may support a cause-and-effect conclusion about treatments. Random sampling supports generalizing to the population represented by the sample. The t procedure alone does not establish either feature of the design.

A strong written justification connects the procedure to the parameter and design, then addresses relevant conditions. For example: “Use a paired t interval for \(\mu_d\), where \(d\) is old minus new energy use, because each randomly selected machine was measured under both settings. The random sample, 10% condition, independent machines, and roughly symmetric differences with no apparent outliers support the procedure.” That explanation makes clear both why the method fits and what quantity it estimates.

Key takeaway: To avoid procedure-selection errors, audit the response, parameter, design link, and question goal in that order. A benchmark is not a second sample, equal sample sizes do not create pairs, and the wording of the goal determines whether the matching t procedure is an interval or a test.

Check Your Understanding

For each situation, identify the target and the procedure, and state the design clue that matters most.

  1. A random sample of 25 packages is used to test whether the population mean mass differs from a fixed label value. Is this one-sample, paired, or two-sample t inference?
  2. The same 12 runners complete a course before and after a training plan. A researcher wants to estimate the mean change in completion time. What parameter and procedure fit?
  3. Two groups each contain 18 randomly assigned students, and every student uses only one of two study methods. Does equal group size make the data paired? Explain.
  4. A researcher asks whether a population mean is greater than a benchmark. Should the result be a confidence interval or a test? What clue in the question decides?
  5. Why can calculator output not establish that a two-sample t procedure is appropriate for two columns of data?