From a Mean Question to a Population Parameter
Once you know that a question concerns a quantitative response, the next task is to say exactly which population quantity the study is about. “The average” is not yet a complete parameter definition. You must identify whose measurements are being averaged, what is measured, and—if there is a comparison—how the means are arranged.
A parameter describes a population. A sample mean, such as \(\bar{x}\), describes the observed sample and is a statistic. In an inference question, the parameter is usually unknown; the sample statistic provides information about it. As in “Defining Both Population Means in Context,” a good definition specifies the target population, the quantitative variable, and the setting or time period. Here we will focus on deciding whether the target is one mean or a mean difference, and on writing the difference precisely.
Three common targets are a single population mean, a population mean of paired differences, and the difference between two population means. The question’s design and wording help distinguish them. In each case, name the response variable first. Then decide which units contribute values to the target quantity.
A Parameter-Finding Routine
Use this routine before writing hypotheses or choosing an inference procedure. It is a way to clarify the question, not a substitute for checking the study design later.
State what is measured for each person, object, or other observational unit, including its units.
Specify which units the question is about and include the relevant setting or time period.
If the question asks about one overall average, name one population mean. If it compares groups or linked measurements, determine what difference the question asks about.
For linked observations, define each unit’s difference first. For two groups, state which group’s mean comes first. Keep that order throughout.
A useful distinction is the source of the two numbers being compared. If each unit provides two connected measurements—for example, a measurement before and after a program—the analysis can be based on a difference for each unit. If units belong to one of two separate, unlinked groups, the comparison is between two population means. The phrase “two averages” alone does not tell you which situation applies.
One Population Mean
A question about the average value of one quantitative response in one target population concerns a single population mean, written \(\mu\). The definition should say what the response is and which population’s measurements are included. If a sample is used, its sample mean \(\bar{x}\) is an estimate, not the parameter itself.
Worked Example: Battery Runtime From One Production Line
A quality team selects 5 rechargeable batteries from a production line’s output during October and measures each battery’s runtime, in hours, under a specified test. The question asks for the true average runtime of all batteries produced by that line during October. The five observed runtimes are 8.4, 7.9, 8.1, 9.0, and 8.6 hours.
Identify the response and population: The quantitative response is battery runtime in hours. The target population is all batteries produced by this line during October. The question asks for one average, not a comparison between groups or two measurements on the same battery.
Name the parameter: Let \(\mu\) be the true mean runtime, in hours, for all rechargeable batteries produced by this production line during October. This definition describes the population quantity the question asks about.
Separate the sample statistic: The five sample runtimes sum to \(8.4+7.9+8.1+9.0+8.6=42.0\) hours. Their sample mean is
The value 8.4 hours is the sample mean, not the value of \(\mu\). It summarizes the five selected batteries. The parameter remains the true mean for the full October production population and is not known just because the sample mean has been calculated.
Paired Measurements: Define the Difference First
When the same unit is measured twice, or units are deliberately matched, the two observations are linked. A helpful approach is to define one difference \(d\) for each pair, using a consistent subtraction order. The parameter is then \(\mu_d\), the true mean of those differences for the target population of pairs. The sign and interpretation depend on the order you choose.
For example, if \(d=\text{after}-\text{before}\), a negative difference means the after measurement is smaller than the before measurement. If the question asks about an average reduction, \(d=\text{before}-\text{after}\) may make a reduction positive instead. Either order can describe the same paired data, but the definition and any later hypotheses must use the chosen order consistently.
Worked Example: Change in Daily Screen Time
A program asks participants to track their daily recreational screen time before and after using a planning tool. The research question asks whether participants’ average daily screen time is reduced after using the tool. Four participants have the following illustrative measurements, in minutes per day:
| Participant | Before | After |
|---|---|---|
| 1 | 31 | 27 |
| 2 | 35 | 30 |
| 3 | 29 | 28 |
| 4 | 33 | 29 |
Identify the observational unit and response: Each participant contributes two linked measurements of the same quantitative response: daily recreational screen time in minutes. The two values from one participant form a pair; they are not observations from two independent groups.
Choose and apply an order: Because the question describes a reduction, define \(d=\text{before}-\text{after}\), in minutes per day. The four sample differences are \(31-27=4\), \(35-30=5\), \(29-28=1\), and \(33-29=4\). Positive values represent reductions.
Name the parameter: Let \(\mu_d\) be the true mean reduction in daily recreational screen time, in minutes per day, among the target population of participants using this planning tool in the specified program setting. The definition names the mean of the paired reductions, not a separate before mean and after mean.
Distinguish parameter from statistic: The observed differences sum to \(4+5+1+4=14\) minutes per day, so their sample mean is
The 3.5 minutes per day is the sample’s average reduction. It is a statistic used to learn about \(\mu_d\); it does not by itself establish the exact population mean reduction.
This distinction is central: paired observations are transformed into one difference per pair, and the parameter is the mean of those differences. In contrast, a comparison of independent groups uses a difference between two population means. The setup of the observations, not merely the fact that the question uses the word “difference,” determines how to name the parameter.
Two Independent Groups: Order the Population Means
When the question compares a quantitative response between two distinct, unlinked populations, define a mean for each population. Then state the difference in a fixed order, such as group 1 minus group 2. As in “Defining Both Population Means in Context,” each mean definition should describe the same response variable and identify its own group and target setting.
The order matters because \(\mu_1-\mu_2\) and \(\mu_2-\mu_1\) have opposite signs. If group 1 is listed first, a positive difference means its population mean is greater than group 2’s. A negative difference means it is smaller. The order is a choice, but it must be stated and kept consistent with the research question.
Worked Example: Delivery Times on Two Routes
A delivery company wants to compare the mean delivery time for completed weekday deliveries on Route North and Route South in June. Different deliveries are assigned to the two routes; no delivery is measured on both routes. A small illustrative sample has times of 42, 46, 44, and 48 minutes for Route North, and 39, 41, 43, and 45 minutes for Route South.
Identify the response and groups: Delivery time in minutes is the quantitative response. The two populations are weekday deliveries completed on Route North in June and weekday deliveries completed on Route South in June. The deliveries are in separate, unlinked groups.
Define the population means and their order: Let \(\mu_N\) be the true mean delivery time, in minutes, for all weekday deliveries completed on Route North in June. Let \(\mu_S\) be the corresponding true mean delivery time for Route South. Use North minus South as the comparison parameter, \(\mu_N-\mu_S\).
Use the sample only to illustrate the distinction: The Route North sample mean is \((42+46+44+48)/4=180/4=45\) minutes. The Route South sample mean is \((39+41+43+45)/4=168/4=42\) minutes. Their observed difference, North minus South, is \(45-42=3\) minutes.
The value 3 minutes is the observed difference in sample means, \(\bar{x}_N-\bar{x}_S\). The parameter is \(\mu_N-\mu_S\), the unknown difference between the two population mean delivery times. With this order, a positive parameter value would mean that the population mean delivery time is greater on Route North; a negative value would mean it is greater on Route South.
Read the Question’s Target Carefully
A scenario may mention several quantities, but only one may be the target of the research question. For example, a question about the average change asks about \(\mu_d\), while a question asking whether the average after-program value differs from a stated benchmark may concern a single population mean. Similarly, a study can record two group labels and several measurements, but you should define the parameter for the response and comparison actually named in the question.
The population in the definition must also match the scope of the question. “All customers” may be too broad if the data concern only customers who placed an online order in a particular month. Include a setting or time period when it distinguishes the target population. If the scenario does not give enough information to identify the population, do not invent a broader claim; state the population the question or design supports.
Common Mistakes and AP Exam Tips
- Writing a sample statistic as the parameter: \(\bar{x}\) and \(\bar{d}\) describe sample data; \(\mu\) and \(\mu_d\) describe population quantities. A full-credit definition names the population parameter, even when a sample mean is supplied.
- Leaving out what is averaged: Writing only “Let \(\mu\) be the mean” does not identify the variable, population, or units. State, for example, “the true mean delivery time, in minutes, for weekday deliveries completed on Route North in June.”
- Calling paired data two independent populations: If each participant contributes both measurements, define a difference for each participant and name \(\mu_d\). Do not describe the pair as if its two measurements came from unrelated groups.
- Failing to define the paired subtraction order: “Mean change” can have either sign unless change is defined. Write \(d=\text{after}-\text{before}\), or another explicit order, and interpret the sign accordingly.
- Leaving a two-group difference unordered: “The difference in mean time” does not say which mean is subtracted from which. Name \(\mu_1-\mu_2\) and identify group 1 and group 2 in context.
- Mixing parameter and sample difference: In a two-group example, \(\bar{x}_1-\bar{x}_2\) is calculated from the samples; \(\mu_1-\mu_2\) is the population comparison. Explain which one the question asks about.
- Choosing a parameter from the word “average” alone: An average can refer to one population, paired changes, or a comparison between group means. Check how the observational units and measurements are organized before naming the target.
For a concise AP response, write a complete parameter sentence before moving to the procedure. For a single population, identify the true mean response in the target population. For paired data, define \(d\) and then identify \(\mu_d\). For independent groups, define both population means and state their difference in a consistent order. The next decision—whether the setting calls for a one-sample, paired, or two-sample procedure—builds on this correctly named target.
Check Your Understanding
For each situation, identify the target parameter and define it in context. When a difference is involved, state the order clearly.
- A city samples household water bills from one neighborhood for April and asks for the true average bill, in dollars, for all households in that neighborhood that month. What is the parameter?
- A trainer records each runner’s 5-kilometer time before and after a training plan and asks for the mean improvement in seconds. Define a paired difference and name the parameter.
- A library compares the mean time, in minutes, that visitors spend in its downtown and east-side branches. Define both population means and a comparison parameter using downtown minus east-side.
- A sample of 30 plants has an average height of 18 centimeters. Explain why 18 centimeters is not automatically the population parameter, and state what \(\mu\) could represent if the target is all plants in the greenhouse.
- A researcher compares heart rates measured on the same participants before and after a short exercise session. Is the target naturally a difference between independent population means or a mean of paired differences? Explain what must be defined first.