Tutorials › AP Statistics › Identifying the Parameter in a Mean Problem

Choosing a mean-inference procedure · Tutorial 762 of 1000

Identifying the Parameter in a Mean Problem

Identify which population mean or mean difference a scenario asks about, and define it clearly before choosing an inference procedure.

Intermediate 9 min read

What You'll Learn

  • Distinguish a population parameter from a sample mean or sample difference.
  • Identify when a question concerns one population mean.
  • Define a difference for paired observations before naming its population mean.
  • Distinguish a paired mean difference from a difference between two independent population means.
  • Choose and interpret a consistent order for group differences.
  • Include the population, quantitative variable, setting, and units in a parameter definition.

From a Mean Question to a Population Parameter

Once you know that a question concerns a quantitative response, the next task is to say exactly which population quantity the study is about. “The average” is not yet a complete parameter definition. You must identify whose measurements are being averaged, what is measured, and—if there is a comparison—how the means are arranged.

A parameter describes a population. A sample mean, such as \(\bar{x}\), describes the observed sample and is a statistic. In an inference question, the parameter is usually unknown; the sample statistic provides information about it. As in “Defining Both Population Means in Context,” a good definition specifies the target population, the quantitative variable, and the setting or time period. Here we will focus on deciding whether the target is one mean or a mean difference, and on writing the difference precisely.

Definition: A population parameter for a mean problem is the true mean of a quantitative variable for a clearly identified population, or the true mean difference defined for a population of pairs or for two populations. It is not the observed mean or difference calculated from the sample.

Three common targets are a single population mean, a population mean of paired differences, and the difference between two population means. The question’s design and wording help distinguish them. In each case, name the response variable first. Then decide which units contribute values to the target quantity.

A Parameter-Finding Routine

Use this routine before writing hypotheses or choosing an inference procedure. It is a way to clarify the question, not a substitute for checking the study design later.

1
Name the quantitative response.
State what is measured for each person, object, or other observational unit, including its units.
2
Identify the population or populations.
Specify which units the question is about and include the relevant setting or time period.
3
Look for a comparison.
If the question asks about one overall average, name one population mean. If it compares groups or linked measurements, determine what difference the question asks about.
4
Define the difference and its order.
For linked observations, define each unit’s difference first. For two groups, state which group’s mean comes first. Keep that order throughout.

A useful distinction is the source of the two numbers being compared. If each unit provides two connected measurements—for example, a measurement before and after a program—the analysis can be based on a difference for each unit. If units belong to one of two separate, unlinked groups, the comparison is between two population means. The phrase “two averages” alone does not tell you which situation applies.

One Population Mean

A question about the average value of one quantitative response in one target population concerns a single population mean, written \(\mu\). The definition should say what the response is and which population’s measurements are included. If a sample is used, its sample mean \(\bar{x}\) is an estimate, not the parameter itself.

Parameter form: Let \(\mu\) be the true mean [quantitative variable, with units] for all [identified observational units] in [the target setting or time period].

Worked Example: Battery Runtime From One Production Line

A quality team selects 5 rechargeable batteries from a production line’s output during October and measures each battery’s runtime, in hours, under a specified test. The question asks for the true average runtime of all batteries produced by that line during October. The five observed runtimes are 8.4, 7.9, 8.1, 9.0, and 8.6 hours.

Identify the response and population: The quantitative response is battery runtime in hours. The target population is all batteries produced by this line during October. The question asks for one average, not a comparison between groups or two measurements on the same battery.

Name the parameter: Let \(\mu\) be the true mean runtime, in hours, for all rechargeable batteries produced by this production line during October. This definition describes the population quantity the question asks about.

Separate the sample statistic: The five sample runtimes sum to \(8.4+7.9+8.1+9.0+8.6=42.0\) hours. Their sample mean is

$$ \bar{x}=\frac{42.0}{5}=8.4\text{ hours} $$

The value 8.4 hours is the sample mean, not the value of \(\mu\). It summarizes the five selected batteries. The parameter remains the true mean for the full October production population and is not known just because the sample mean has been calculated.

Paired Measurements: Define the Difference First

When the same unit is measured twice, or units are deliberately matched, the two observations are linked. A helpful approach is to define one difference \(d\) for each pair, using a consistent subtraction order. The parameter is then \(\mu_d\), the true mean of those differences for the target population of pairs. The sign and interpretation depend on the order you choose.

For example, if \(d=\text{after}-\text{before}\), a negative difference means the after measurement is smaller than the before measurement. If the question asks about an average reduction, \(d=\text{before}-\text{after}\) may make a reduction positive instead. Either order can describe the same paired data, but the definition and any later hypotheses must use the chosen order consistently.

Definition: For paired observations, \(d\) is the difference between the two measurements on one pair, defined in a stated order. The parameter \(\mu_d\) is the true mean of those pairwise differences for the target population of pairs.

Worked Example: Change in Daily Screen Time

A program asks participants to track their daily recreational screen time before and after using a planning tool. The research question asks whether participants’ average daily screen time is reduced after using the tool. Four participants have the following illustrative measurements, in minutes per day:

ParticipantBeforeAfter
13127
23530
32928
43329

Identify the observational unit and response: Each participant contributes two linked measurements of the same quantitative response: daily recreational screen time in minutes. The two values from one participant form a pair; they are not observations from two independent groups.

Choose and apply an order: Because the question describes a reduction, define \(d=\text{before}-\text{after}\), in minutes per day. The four sample differences are \(31-27=4\), \(35-30=5\), \(29-28=1\), and \(33-29=4\). Positive values represent reductions.

Name the parameter: Let \(\mu_d\) be the true mean reduction in daily recreational screen time, in minutes per day, among the target population of participants using this planning tool in the specified program setting. The definition names the mean of the paired reductions, not a separate before mean and after mean.

Distinguish parameter from statistic: The observed differences sum to \(4+5+1+4=14\) minutes per day, so their sample mean is

$$ \bar{d}=\frac{14}{4}=3.5\text{ minutes per day} $$

The 3.5 minutes per day is the sample’s average reduction. It is a statistic used to learn about \(\mu_d\); it does not by itself establish the exact population mean reduction.

This distinction is central: paired observations are transformed into one difference per pair, and the parameter is the mean of those differences. In contrast, a comparison of independent groups uses a difference between two population means. The setup of the observations, not merely the fact that the question uses the word “difference,” determines how to name the parameter.

Two Independent Groups: Order the Population Means

When the question compares a quantitative response between two distinct, unlinked populations, define a mean for each population. Then state the difference in a fixed order, such as group 1 minus group 2. As in “Defining Both Population Means in Context,” each mean definition should describe the same response variable and identify its own group and target setting.

Parameter form: Let \(\mu_1\) and \(\mu_2\) be the true means of the same quantitative variable for the two identified populations. The comparison parameter is \(\mu_1-\mu_2\), in the units of that variable.

The order matters because \(\mu_1-\mu_2\) and \(\mu_2-\mu_1\) have opposite signs. If group 1 is listed first, a positive difference means its population mean is greater than group 2’s. A negative difference means it is smaller. The order is a choice, but it must be stated and kept consistent with the research question.

Worked Example: Delivery Times on Two Routes

A delivery company wants to compare the mean delivery time for completed weekday deliveries on Route North and Route South in June. Different deliveries are assigned to the two routes; no delivery is measured on both routes. A small illustrative sample has times of 42, 46, 44, and 48 minutes for Route North, and 39, 41, 43, and 45 minutes for Route South.

Identify the response and groups: Delivery time in minutes is the quantitative response. The two populations are weekday deliveries completed on Route North in June and weekday deliveries completed on Route South in June. The deliveries are in separate, unlinked groups.

Define the population means and their order: Let \(\mu_N\) be the true mean delivery time, in minutes, for all weekday deliveries completed on Route North in June. Let \(\mu_S\) be the corresponding true mean delivery time for Route South. Use North minus South as the comparison parameter, \(\mu_N-\mu_S\).

Use the sample only to illustrate the distinction: The Route North sample mean is \((42+46+44+48)/4=180/4=45\) minutes. The Route South sample mean is \((39+41+43+45)/4=168/4=42\) minutes. Their observed difference, North minus South, is \(45-42=3\) minutes.

The value 3 minutes is the observed difference in sample means, \(\bar{x}_N-\bar{x}_S\). The parameter is \(\mu_N-\mu_S\), the unknown difference between the two population mean delivery times. With this order, a positive parameter value would mean that the population mean delivery time is greater on Route North; a negative value would mean it is greater on Route South.

Read the Question’s Target Carefully

A scenario may mention several quantities, but only one may be the target of the research question. For example, a question about the average change asks about \(\mu_d\), while a question asking whether the average after-program value differs from a stated benchmark may concern a single population mean. Similarly, a study can record two group labels and several measurements, but you should define the parameter for the response and comparison actually named in the question.

The population in the definition must also match the scope of the question. “All customers” may be too broad if the data concern only customers who placed an online order in a particular month. Include a setting or time period when it distinguishes the target population. If the scenario does not give enough information to identify the population, do not invent a broader claim; state the population the question or design supports.

Key takeaway: First name the quantitative response and the units. Then identify whether the target is one population mean \(\mu\), a mean of paired differences \(\mu_d\), or a difference between two population means such as \(\mu_1-\mu_2\). Define the population and subtraction order in context; keep sample statistics separate from parameters.

Common Mistakes and AP Exam Tips

  • Writing a sample statistic as the parameter: \(\bar{x}\) and \(\bar{d}\) describe sample data; \(\mu\) and \(\mu_d\) describe population quantities. A full-credit definition names the population parameter, even when a sample mean is supplied.
  • Leaving out what is averaged: Writing only “Let \(\mu\) be the mean” does not identify the variable, population, or units. State, for example, “the true mean delivery time, in minutes, for weekday deliveries completed on Route North in June.”
  • Calling paired data two independent populations: If each participant contributes both measurements, define a difference for each participant and name \(\mu_d\). Do not describe the pair as if its two measurements came from unrelated groups.
  • Failing to define the paired subtraction order: “Mean change” can have either sign unless change is defined. Write \(d=\text{after}-\text{before}\), or another explicit order, and interpret the sign accordingly.
  • Leaving a two-group difference unordered: “The difference in mean time” does not say which mean is subtracted from which. Name \(\mu_1-\mu_2\) and identify group 1 and group 2 in context.
  • Mixing parameter and sample difference: In a two-group example, \(\bar{x}_1-\bar{x}_2\) is calculated from the samples; \(\mu_1-\mu_2\) is the population comparison. Explain which one the question asks about.
  • Choosing a parameter from the word “average” alone: An average can refer to one population, paired changes, or a comparison between group means. Check how the observational units and measurements are organized before naming the target.

For a concise AP response, write a complete parameter sentence before moving to the procedure. For a single population, identify the true mean response in the target population. For paired data, define \(d\) and then identify \(\mu_d\). For independent groups, define both population means and state their difference in a consistent order. The next decision—whether the setting calls for a one-sample, paired, or two-sample procedure—builds on this correctly named target.

Check Your Understanding

For each situation, identify the target parameter and define it in context. When a difference is involved, state the order clearly.

  1. A city samples household water bills from one neighborhood for April and asks for the true average bill, in dollars, for all households in that neighborhood that month. What is the parameter?
  2. A trainer records each runner’s 5-kilometer time before and after a training plan and asks for the mean improvement in seconds. Define a paired difference and name the parameter.
  3. A library compares the mean time, in minutes, that visitors spend in its downtown and east-side branches. Define both population means and a comparison parameter using downtown minus east-side.
  4. A sample of 30 plants has an average height of 18 centimeters. Explain why 18 centimeters is not automatically the population parameter, and state what \(\mu\) could represent if the target is all plants in the greenhouse.
  5. A researcher compares heart rates measured on the same participants before and after a short exercise session. Is the target naturally a difference between independent population means or a mean of paired differences? Explain what must be defined first.