From a Test Decision to a Complete Conclusion
A mean test ends with more than “reject” or “fail to reject.” The final conclusion should answer the research question in context: what population mean or difference in means is being studied, what direction the evidence supports, and whether the data provide convincing evidence for the claim. In “Tying the Conclusion to the Original Claim” and “Common Mistakes in Writing Conclusions,” you practiced these ideas. Here, you will apply them across one-sample, paired, and two-sample settings.
The core wording stays consistent across these settings, but the parameter changes. A one-sample test concerns one population mean, \(\mu\). A paired test concerns the population mean of the pairwise differences, \(\mu_d\). A two-sample test concerns a difference between two population means, such as \(\mu_1-\mu_2\). Make that parameter clear in the conclusion rather than relying on a vague phrase such as “the average is different.”
A Reliable Shape for the Final Sentence
As in “Making a Decision From P-Value and Alpha,” reject \(H_0\) when \(p\le\alpha\), and fail to reject \(H_0\) when \(p>\alpha\). The evidence statement follows that decision: rejection supports saying the data provide convincing evidence for the alternative; failure to reject calls for saying they do not provide convincing evidence for the alternative. Neither decision proves a claim.
A useful structure is: “Since \(p\) [comparison] \(\alpha\), we [decision]. The data [do/do not] provide convincing evidence that [contextual alternative claim about the population parameter].” You may combine the two parts into one sentence, but both should be clear. Include the numerical p-value and \(\alpha\) when they are provided or requested.
- Compare the p-value with the stated significance level and give the correct decision.
- Name the population mean or difference in means, not just the sample statistic.
- Describe the variable and population in the setting.
- Match the claim’s direction to \(H_a\), including the order of groups or definition of paired differences.
- Do not claim proof, certainty, or equality. Do not generalize or make a causal claim beyond what the design supports.
The final point matters because statistical evidence and study design answer different questions. As explained in “Generalizing Conclusions to a Population,” sampling determines what population a result may represent. Random assignment can support a causal interpretation of treatment differences; random sampling can support generalization to a population. A p-value alone provides neither kind of support.
One-Sample Mean: Name the Population and Benchmark
For a one-sample test, the conclusion should identify the population whose mean is being tested and the benchmark in the alternative claim. Do not substitute the sample mean for the population mean. The sample mean is evidence used in the test; the conclusion concerns the population mean.
Worked Example: Mean Fill Volume in a Bottling Run
A quality team takes a random sample of bottles from a particular day’s production run and tests whether the true mean fill volume is below the labeled 500 milliliters. A one-sample t test gives \(p=0.018\), and the team uses \(\alpha=0.05\). The test is appropriate for this setting, and the question asks for the final conclusion.
Decision: Compare the supplied values: \(0.018<0.05\). Because \(p<\alpha\), reject \(H_0\).
Parameter and claim: The parameter is the true mean fill volume, in milliliters, for bottles in this production run. The alternative claim is that this mean is less than 500 milliliters. The conclusion should not say that every bottle is underfilled; the test concerns a population mean.
Complete conclusion: “Since \(p=0.018<0.05\), we reject \(H_0\). The data provide convincing evidence that the true mean fill volume of bottles in this day’s production run is less than 500 milliliters.”
This conclusion gives the decision, evidence wording, population, variable with units, and direction of the claim. It does not state that the sample mean is the true mean or that every bottle has less than 500 milliliters.
Paired Mean: Make the Difference’s Sign Meaningful
A paired test is a one-sample t test applied to pairwise differences. As in “Two-Sample Versus Paired Test on the Same Numbers,” the data’s genuine pairing determines the procedure. A conclusion should state what a positive or negative difference means in the setting. Define the subtraction order consistently with the test’s alternative; otherwise, a correct decision can be followed by a claim in the wrong direction.
Worked Example: Comparing Two Scheduling Systems for Dispatchers
A company randomly selects 24 dispatchers. Each dispatcher completes comparable scheduling tasks using both the usual system and a new system, with the order randomly assigned. Let \(d=\text{usual-system time}-\text{new-system time}\), measured in minutes. A paired t test evaluates whether the true mean difference \(\mu_d\) is greater than 0; this would mean that the new system has a lower mean task-completion time. The test gives \(p=0.031\), with \(\alpha=0.05\).
Decision: \(0.031<0.05\), so reject \(H_0\).
Interpret the parameter: Here \(\mu_d\) is the true mean of the dispatcher-level differences, usual-system time minus new-system time, for dispatchers in the company. A positive mean difference corresponds to less time with the new system. Keep the subtraction order visible when translating the alternative into words.
Complete conclusion: “Since \(p=0.031<0.05\), we reject \(H_0\). The data provide convincing evidence that, for dispatchers at this company, the true mean task-completion time is lower with the new scheduling system than with the usual system.”
The random sample supports generalizing to the company’s dispatchers, and the randomized order helps address treatment-order effects. The conclusion is about a mean difference, not a guarantee that the new system saves time for every dispatcher or every task.
If the subtraction order were reversed, the numerical differences and alternative direction would also need to be reversed. For example, defining \(d=\text{new-system time}-\text{usual-system time}\) would make evidence for a lower mean time with the new system correspond to \(\mu_d<0\), not \(\mu_d>0\). Always interpret the defined difference rather than memorizing “greater means better” or “less means worse.”
Two-Sample Mean: Keep the Group Order in the Conclusion
For two independent groups, define the difference in the same order used by the hypotheses. If the parameter is \(\mu_1-\mu_2\), a positive difference means group 1’s population mean is greater than group 2’s, and a negative difference means it is lower. The conclusion must preserve this order. The full test structure and parameter definitions were covered in “Writing the Full Two-Sample t Test Solution” and “Defining Both Population Means in Context.”
Worked Example: Comparing Water-Quality Measurements at Two Ponds
Environmental students independently take random samples of water from Pond North and Pond South in a local park. They compare the mean nitrate concentration, in milligrams per liter, for the two ponds. Let \(\mu_N-\mu_S\) be the difference between the true mean concentrations at North and South. A two-sided two-sample t test gives \(p=0.14\), with \(\alpha=0.05\).
Decision: \(0.14>0.05\), so fail to reject \(H_0\).
Interpret the alternative: The two-sided alternative asks whether the population mean concentrations differ in either direction. Since the test does not reject, the correct conclusion is not that the means are equal. It is that the data do not provide convincing evidence of a difference in the two population means.
Complete conclusion: “Since \(p=0.14>0.05\), we fail to reject \(H_0\). The data do not provide convincing evidence that the true mean nitrate concentration differs between water in Pond North and water in Pond South in this park.”
The conclusion gives the two populations and the measured variable, and it respects the two-sided alternative. Because the samples were randomly selected from these ponds, the comparison can describe their water populations under the sampling design. The result does not establish that the mean concentrations are equal.
When the Evidence Is Not Convincing
A fail-to-reject conclusion is not a weaker version of a rejection conclusion. It answers the question carefully: the observed data do not provide convincing evidence for the alternative claim at the chosen significance level. It does not say that the null hypothesis is true, that the means are identical, or that an effect is impossible. As discussed in “Why a t Test Never Proves the Null Mean,” a test may fail to detect a difference even when one exists.
For example, suppose a paired test asks whether a new stretching routine reduces athletes’ mean recovery time, and \(p=0.12\) at \(\alpha=0.05\). The comparison is \(0.12>0.05\), so the decision is to fail to reject \(H_0\). A suitable conclusion would say that the data do not provide convincing evidence that the routine reduces the population mean recovery time for the athletes represented by the study. It would not say “the routine has no effect.”
The same evidence wording applies whether the alternative is one-sided or two-sided, but the claim must match that alternative. For a one-sided test, name the specified direction. For a two-sided test, say “differs” or “is different,” not “is higher” unless the test was actually designed to support that directional claim.
Common Mistakes and AP Exam Tips
- Stopping at the decision: “Reject \(H_0\)” does not answer the research question. Add a contextual statement about what the data provide convincing evidence for.
- Using the sample result as the conclusion: “The sample mean is lower” reports a statistic, not the population claim. A full-credit conclusion names the true population mean or mean difference.
- Reversing a paired difference: If \(d\) is defined as before minus after, interpret its sign using that order. State what a positive or negative difference means before writing the claim.
- Reversing two groups: For \(\mu_1-\mu_2\), keep group 1 first and group 2 second. “Group 2 is lower” may be mathematically equivalent to “group 1 is higher,” but unclear switching makes the direction hard to grade.
- Claiming equality after failing to reject: A large p-value is not evidence that two population means are exactly equal. Say the data do not provide convincing evidence for the alternative.
- Overstating what the design supports: Do not claim causation from an observational comparison, or generalize to a population not represented by the sampling method. Avoid “proves,” “always,” and “for everyone.”
A quick final check is to read your conclusion without looking at the hypotheses. Can a reader tell which population or groups are being discussed, what quantitative variable is measured, and whether the evidence supports a higher, lower, or different mean? If not, revise the contextual claim. Then check the study description: does it justify the population scope and any causal wording you used?
Check Your Understanding
For each situation, write a complete final conclusion. Include the decision and a contextual statement about the evidence.
- A one-sample test asks whether the true mean battery life for a model of tablet is greater than 9 hours. The p-value is 0.022 and \(\alpha=0.05\). What should the conclusion say?
- A paired test defines \(d=\text{week 1 score}-\text{week 4 score}\) for the same students and tests whether \(\mu_d<0\). The p-value is 0.19 and \(\alpha=0.10\). What does the sign convention mean, and what conclusion follows?
- A two-sample test uses \(\mu_A-\mu_B\) for the true mean commute times of two neighborhoods and has a two-sided alternative. The p-value is 0.008 and \(\alpha=0.01\). Write a conclusion without claiming which neighborhood has the greater mean.
- Why is “there is no difference between the two population means” not an appropriate conclusion when a two-sided test fails to reject \(H_0\)?
- What design information should you check before a conclusion generalizes to a population or claims that a treatment caused a change?