First Identify the Goal
In “Spotting Independent Samples in Word Problems,” you looked at how a study’s design can tell you whether two groups are paired or unpaired. Now consider a different first question: What does the problem want you to find out? A mean-inference question may ask you to estimate a population parameter, or it may ask whether the data provide evidence for or against a claim about that parameter.
Those goals point to different forms of inference. An interval gives a range of plausible values for a population parameter, with a stated confidence level. A test evaluates how convincing the sample evidence is against a null hypothesis. The choice is not based just on whether the problem includes a number or uses the word “mean.” Look for the task the question asks you to complete.
This choice comes before selecting the specific mean procedure. As in “Identifying the Parameter in a Mean Problem,” first identify the population parameter and its units. Then use the study design—one sample, paired data, or two independent samples—to choose the appropriate interval or test for that parameter.
Wording That Usually Signals an Interval
Estimation questions often ask how much, what value, or within what range. Phrases such as “estimate the true mean,” “find a plausible range,” “how large is the difference,” and “give a 95% confidence interval” point toward an interval. The interval does not claim that every value in the range is equally likely or that the parameter changes from sample to sample. It uses sample data to estimate a fixed population parameter.
A point estimate, such as \(\bar{x}\), is one estimate of a population mean \(\mu\). A confidence interval adds a margin of error to express uncertainty in that estimate. For a mean, the confidence interval’s endpoints are in the response variable’s units. For a difference in means, they are in units of the difference, such as minutes or points.
A question may ask for a confidence interval without stating a particular null value to test. Do not invent a claim merely because the interval might later be compared with one. State the parameter, find the requested interval, and interpret it in context.
Wording That Usually Signals a Test
A testing question asks whether the evidence supports a claim about a population parameter. Watch for wording such as “is there convincing evidence,” “do the data support the claim,” “has the mean increased,” “is the mean different,” or “is the average greater than the benchmark?” These questions call for hypotheses and a test, not just an estimated range.
For a mean, hypotheses concern a population parameter such as \(\mu\), a paired-difference mean \(\mu_d\), or a difference in population means \(\mu_1-\mu_2\). The null hypothesis states a reference claim, often equality to a specified value. The alternative hypothesis expresses the direction or difference the question asks about. As covered in “Writing a Full Two-Sample t Test Solution,” a complete test response connects the parameter and hypotheses to the design and conditions, the calculation, and a contextual conclusion.
A test does not determine the probability that a claim is true. Its p-value describes how unusual the observed test statistic, or a more extreme one in the alternative’s direction, would be if the null hypothesis were true. Use the conclusion language established in “Writing a Conclusion for a Two-Sample t Test”: reject or fail to reject \(H_0\), then state what the data do or do not provide convincing evidence for in context.
Keep the Goal Separate From the Design
The wording tells you whether the goal is estimation or testing. The study design tells you which parameter and procedure fit. Do not confuse these two decisions. For example, a question about the average change for the same people before and after a program could ask either for a confidence interval for \(\mu_d\) or a test about \(\mu_d\). The repeated measurements make the data paired; the question’s goal determines interval versus test.
Likewise, a comparison of two independent groups can lead to a confidence interval for \(\mu_1-\mu_2\) or a test about that difference. As in “Spotting Independent Samples in Word Problems,” the groups’ structure comes from how the units were selected or assigned, not from the wording “estimate” or “evidence.”
Does the question ask for a range or an estimate, or does it ask whether evidence supports a claim?
Identify the population mean or mean difference, including the response and units.
Decide whether the data concern one mean, paired differences, or the difference between two independent means.
Give an interval for estimation; use hypotheses, a test, and a contextual evidence conclusion for a claim-testing task.
Worked Examples
Worked Example: Estimating a Mean Commute Time
A transit planner takes a random sample of 16 bus commuters from a city population of 1,200 commuters. The sample mean commute time is 28 minutes, and the sample standard deviation is 8 minutes. The planner asks, “What is a reasonable 95% confidence interval for the true mean commute time?”
Identify the goal: “What is a reasonable ... confidence interval?” explicitly asks for an estimate of a population mean. This is an interval task, not a test of a claim. Let \(\mu\) be the true mean commute time, in minutes, for the city’s bus commuters.
Choose the procedure and check conditions: The data are one random sample, so a one-sample t interval for \(\mu\) is the candidate procedure. The sample is less than 10% of the population because \(16/1200=0.0133\), or about 1.3%, so the 10% condition is satisfied. Suppose an inspection of the sample’s distribution shows no strong skewness or outliers; with the small sample, that supports using t inference.
Calculate the interval: With \(n=16\), there are \(15\) degrees of freedom. The 95% critical value is \(t^*=2.131\). The standard error is \(s/\sqrt{n}=8/\sqrt{16}=2\) minutes, so the margin of error is \(2.131(2)=4.262\) minutes.
Interpret in context: We are 95% confident that the true mean commute time for the city’s bus commuters is between about 23.7 and 32.3 minutes. The question did not ask whether the mean differs from a benchmark, so a test conclusion would not answer its stated goal.
Worked Example: Testing a Claim About Battery Life
A manufacturer advertises that a model of rechargeable battery lasts an average of 48 hours under a specified test. A random sample of 25 batteries has a mean lifetime of 52 hours and a standard deviation of 10 hours. Assume the population distribution is reasonably compatible with t inference. The question asks, “Do the data provide convincing evidence that the true mean lifetime exceeds 48 hours?”
State: Let \(\mu\) be the true mean lifetime, in hours, of batteries of this model under the specified test. The question asks for evidence that the mean exceeds the advertised value:
Plan: This is a claim-testing task because it asks whether there is convincing evidence for a directional claim. Use a one-sample t test. The sample is random. If the battery population is large enough that 25 is less than 10% of it, the 10% condition is met. The problem states that the population distribution is reasonably compatible with t inference, so the distribution condition is met.
Do: The standard error is \(s/\sqrt{n}=10/\sqrt{25}=2\) hours. The test statistic is \(t=(52-48)/2=2.00\), with \(24\) degrees of freedom. For the right-tailed alternative, the p-value is about \(0.0285\), rounded.
Conclude: At \(\alpha=0.05\), \(0.0285<0.05\), so reject \(H_0\). The data provide convincing evidence that the true mean battery lifetime under the specified test exceeds 48 hours. A confidence interval could estimate the mean, but the question specifically asks whether evidence supports this claim, so the test addresses the stated goal.
Worked Example: Estimating or Testing a Difference Between Groups
A school randomly assigns 40 students to one of two review methods, with 20 students in each group. Each student completes one quiz, and the response is the score in points. Let \(\mu_A\) be the true mean score under Method A and \(\mu_B\) the true mean under Method B. The observed sample means are 82 points and 77 points, respectively.
First version of the question: “How large is the true difference in mean quiz scores between the two methods?” This asks for estimation. The target is \(\mu_A-\mu_B\), in points, so the appropriate form is a two-sample t interval for the difference in means, provided the procedure’s conditions are met. The observed sample difference is \(82-77=5\) points; it is a statistic, not the true difference.
Second version of the question: “Do the data provide convincing evidence that Method A produces a higher mean quiz score than Method B?” This asks for a test. The target remains \(\mu_A-\mu_B\), but the hypotheses now match the directional claim:
Because the students were randomly assigned to one method each and no matching is described, the groups are unpaired. A two-sample t test is the candidate procedure, subject to its conditions. The sample means alone do not provide enough information to calculate its standard error, test statistic, or p-value; the group standard deviations are also needed. The wording still identifies the goal correctly: the first version asks for an interval, and the second asks for a test.
Worked Example: Paired Measurements and Two Possible Goals
A fitness coach records the number of minutes 12 athletes can run before stopping, once before and once after a training plan. Define each athlete’s difference as after minus before. The sample mean difference is 3.5 minutes. Suppose the design and distribution of the differences support paired t inference.
If the question asks for a typical amount of change: “Estimate the true mean increase in running time for athletes like these” is an estimation task. Define \(\mu_d\) as the true mean of the after-minus-before differences, in minutes. The appropriate form is a paired t interval for \(\mu_d\). The observed mean difference, 3.5 minutes, is the point estimate; an interval would quantify uncertainty around it. A numerical interval cannot be calculated from the information given because the standard deviation of the differences is not provided.
If the question asks whether the plan increases mean running time: “Do the data provide convincing evidence of an increase?” is a testing task. The parameter and difference order stay the same, but the hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\). The appropriate form is a paired t test, using the sample of 12 differences. Here, the same athletes were measured twice, so treating the two sets of times as independent groups would ignore the pairing.
The numbers describe the observed change, but they do not determine whether to use an interval or test. The question’s goal does.
Questions That Ask for Both
Some prompts deliberately request more than one result. For example, “Construct a 95% confidence interval for the mean difference, then use it to assess whether the data are consistent with no difference” asks for an interval and an interpretation related to a claim. In that case, do both parts. Do not assume that every question must be answered with only one output.
A confidence interval can also help you understand the size and precision of an estimated effect, while a test focuses on evidence against a null hypothesis. As covered in “Linking Two-Sample Intervals and Tests” and “Connecting Confidence Intervals to Test Decisions,” a two-sided test and a matching confidence interval are connected. That relationship does not mean the goals are identical: follow what the question requests, and be clear about whether you are estimating a parameter or evaluating evidence.
Be especially careful when the claim is that a mean equals a value. A test that fails to reject the null hypothesis does not prove equality. As explained in “Why a t Test Never Proves the Null Mean,” a confidence interval can show which values remain plausible, but neither procedure turns a failure to find convincing evidence into proof of an exact population value.
Common Mistakes and AP Exam Tips
- Choosing a test just because a number appears: A benchmark such as 48 hours matters to a test only when the question asks whether the data provide evidence about a claim involving that benchmark. “Estimate the mean lifetime” still calls for an interval.
- Giving an interval when the question asks for evidence: An interval by itself may not answer “Do the data provide convincing evidence that the mean increased?” State and assess the hypotheses for a test when that is the requested task.
- Letting “significant” replace the question’s goal: If a prompt asks how large a difference is, report an interval that estimates the difference. Do not substitute a test decision for the requested estimate.
- Confusing the parameter with the statistic: The sample mean or sample difference is not the population parameter. Define \(\mu\), \(\mu_d\), or \(\mu_1-\mu_2\) in context and give its units.
- Choosing the goal from the design: Paired data do not automatically mean “test,” and independent samples do not automatically mean “interval.” The design selects the suitable version of the procedure; the question’s wording selects interval versus test.
- Claiming a test proves equality: If you fail to reject \(H_0\), say the data do not provide convincing evidence for the alternative claim. Do not say that the population mean is equal to the null value.
- Ignoring an explicit request for both: If the prompt requests an interval and a test-related assessment, answer both parts rather than choosing one and leaving the other unanswered.
A strong first sentence can make the choice clear: “Because the question asks for a plausible range for the population mean, I will use a confidence interval,” or “Because the question asks whether there is convincing evidence that the population mean exceeds the benchmark, I will conduct a test.” Then identify the parameter and select the procedure that matches the design.
Check Your Understanding
For each prompt, decide whether the main task is estimation, claim testing, or both. Name the population parameter and the appropriate form of inference when the design is specified.
- A random sample of 22 library visitors is used to estimate the true mean time, in minutes, that visitors spend in the building. Is this an interval or a test task?
- A random sample of 30 packages has a mean mass of 2.04 kilograms. The question asks whether the population mean package mass differs from the stated target of 2 kilograms. What is the goal, and what form of inference is called for?
- Researchers measure the same 15 plants’ heights before and after a growing treatment. The prompt asks for a plausible range for the true mean after-minus-before height change. What parameter and form of inference fit?
- Two independent random samples of cyclists are used to investigate whether mean travel time differs between two routes. Does the wording ask for estimation or testing? What is the target parameter?
- A report requests a confidence interval for the mean difference and asks whether zero is a plausible value for that difference. Should you provide an interval, a test, or both requested interpretations?