From a Numerical Result Back to the Claim
A t-test produces numbers, but the research question is usually stated in words: a battery’s mean runtime exceeds a target, a task takes less time after a workshop, or one group has a higher mean than another. A strong conclusion connects those words to the test result. It does not stop at “the result is significant,” and it does not claim more than the data show.
The earlier tutorials “Making a Decision From P-Value and Alpha” and “Writing a Conclusion for a Two-Sample t Test” established how to compare a p-value with \(\alpha\) and choose between rejecting and failing to reject \(H_0\). Here, the focus is the next step: reconciling that decision with the claim that motivated the test. Before writing, check the parameter, the direction of the alternative, the observed sample result, the decision, and the study’s scope.
A claim may not use statistical language. For example, “the new batteries last longer than 12 hours on average” becomes a claim about a population mean: \(\mu>12\) hours. The claim “the average time is different from 12 hours” becomes \(\mu\ne12\). Translating the words into a parameter and a direction helps prevent a conclusion from accidentally answering a different question.
A Claim-to-Conclusion Routine
Use the following routine after identifying the test output. The earlier tutorial “Interpreting Calculator Output for a t Test” explains how to read \(t\), \(df\), and \(p\); the goal here is to put those values into a meaningful conclusion.
Name the parameter, population, measured variable, and units. Check whether the claim is about one mean, a mean difference, or a difference between two means.
Check that the original claim matches \(H_a\). Then check whether the sample result points in the claimed direction. For a two-sample test, pay attention to the order in \(\bar{x}_1-\bar{x}_2\).
Compare \(p\) with the stated \(\alpha\). Reject \(H_0\) when \(p\le\alpha\); otherwise, fail to reject \(H_0\). The sample’s direction alone does not determine this decision.
After rejecting \(H_0\), say the data provide convincing evidence for the claim represented by \(H_a\). After failing to reject, say the data do not provide convincing evidence for that claim. Name the population parameter and keep the conclusion within the study’s scope.
The alternative hypothesis determines what claim a test can support. A test with \(H_a:\mu>12\) can provide evidence that the population mean exceeds 12; it cannot provide evidence that the mean is exactly 12. Likewise, failure to reject \(H_0\) does not establish that the null value is true, as explained in “Why a t Test Never Proves the Null Mean.”
Also distinguish a sample’s direction from the strength of evidence. A sample mean can be above the null value, yet the p-value can be larger than \(\alpha\). In that case, the sample points in the claim’s direction, but the test does not provide convincing evidence for the claim at the chosen significance level.
Worked Example: Does the Mean Runtime Exceed the Target?
Worked Example: Does the Mean Runtime Exceed the Target?
A fictional equipment maker randomly selects 16 rechargeable lantern batteries from a production batch of at least 160 batteries. It records each battery’s runtime, in hours, under the same test conditions. The sample mean is 13.2 hours and the sample standard deviation is 2.4 hours. A plot of the runtimes shows no severe skewness or extreme outliers. The maker wants to know whether the batch’s mean runtime exceeds 12 hours and uses \(\alpha=0.05\).
Let \(\mu\) be the true mean runtime, in hours, of rechargeable lantern batteries in this production batch. The original claim is \(\mu>12\) hours, so the hypotheses are \(H_0:\mu=12\) hours and \(H_a:\mu>12\) hours.
Use a one-sample t test for the batch’s population mean. The batteries were randomly selected. The 10% condition is met because the sample of 16 is no more than 10% of the batch: \(16\le0.10(160)=16\). The plot shows no severe skewness or extreme outliers, supporting a t procedure for this sample size.
The estimated standard error is \(s/\sqrt{n}\). Calculate the test statistic and the right-tailed p-value:
The data provide convincing evidence that the true mean runtime of rechargeable lantern batteries in this production batch exceeds 12 hours. This conclusion answers the original claim and names the population mean it concerns.
The sample mean is above 12 hours, so its direction agrees with the claim. The p-value comparison supplies the separate evidence decision. Saying only “the batteries last longer” would be too vague: it would leave unclear whether the statement refers to these sampled batteries, the batch’s population mean, or every individual battery.
Worked Example: Is Mean Task Time Lower After a Workshop?
Worked Example: Is Mean Task Time Lower After a Workshop?
A fictional community garden randomly selects 9 eligible volunteers from a population of 90. Each volunteer completes the same planting task before and after a short workshop. Define each difference as before time minus after time, in minutes, so a positive difference means the volunteer took less time after the workshop. The mean of the 9 differences is 3.0 minutes, and their standard deviation is 3.0 minutes. A plot of the differences shows no severe skewness or extreme outliers. The garden tests whether the population mean difference is greater than zero, using \(\alpha=0.05\).
Let \(\mu_d\) be the true mean before-minus-after task-time difference, in minutes, for eligible volunteers at this garden. A positive \(\mu_d\) means mean task time is lower after the workshop. Thus, the claim that the population mean task time is lower after the workshop corresponds to \(H_a:\mu_d>0\); the null is \(H_0:\mu_d=0\).
A paired t test is appropriate because each volunteer has two linked measurements, and the analysis uses one difference per volunteer. The volunteers were randomly selected. The 10% condition is met because \(9\le0.10(90)=9\). The plot of the differences supports using a t procedure with this small sample. The paired observations within a volunteer are not treated as independent; the differences are the observations analyzed.
The standard error is \(3.0/\sqrt{9}=1.0\) minute. Therefore,
For a right-tailed t test with \(t=3.000\) and \(df=8\), the p-value is \(0.008536\), or \(0.0085\) rounded to four decimal places. Since \(0.0085<0.05\), reject \(H_0\). The data provide convincing evidence that the true mean before-minus-after task-time difference for eligible volunteers is greater than zero. Equivalently, the data provide convincing evidence that their population mean task time is lower after the workshop.
Notice what the conclusion does—and does not—say. The study supports the stated comparison of mean task times for the population represented by the random sample. Because volunteers were measured before and after rather than randomly assigned to a workshop or a no-workshop group, this comparison by itself does not establish that the workshop caused the reduction.
Worked Example: Does Group A Have a Higher Mean?
Worked Example: Does Group A Have a Higher Mean?
In a fictional comparison, researchers randomly select 16 members of Group A and 16 members of Group B from their respective populations. They measure the same quantitative response for both groups. Group A has a sample mean of 34 units and a sample standard deviation of 8 units; Group B has a sample mean of 34 units and a sample standard deviation of 8 units. Assume the samples are independent, each population contains at least 160 members, and the plots for both groups show no severe skewness or extreme outliers. The research claim is that Group A’s population mean is higher. Use \(\alpha=0.05\).
Let \(\mu_A\) and \(\mu_B\) be the true mean responses, in units, for the populations represented by Groups A and B. The claim gives \(H_a:\mu_A-\mu_B>0\), with \(H_0:\mu_A-\mu_B=0\). An unpooled two-sample t test is appropriate. The groups are independent by the sampling design, and the observations within each random sample are independent. The 10% condition is met for each sample because \(16\le0.10(160)=16\). The plots support using the two-sample t procedure.
The observed difference in sample means is \(34-34=0\) units. The standard error and test statistic are
With equal sample sizes and standard deviations, the unpooled degrees of freedom are \(df=16+16-2=30\). For the right-tailed alternative and \(t=0\), the p-value is \(0.5000\). Since \(0.5000>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that Group A’s population mean response is higher than Group B’s.
This conclusion is not the same as saying that the population means are equal. The sample means happen to be equal, and the test does not show convincing evidence for the higher-than claim. It does not prove the null hypothesis. The result answers the original directional question without turning “not enough evidence for higher” into “the groups are the same.”
Common Mistakes and AP Exam Tips
- Repeating the decision without answering the claim: “Reject \(H_0\)” is a decision, not a complete contextual conclusion. Add what the data provide convincing evidence about, using the population parameter and the original context.
- Reporting only the sample direction: A sample mean above the null value does not automatically support a greater-than claim. Compare the p-value with \(\alpha\) before deciding whether the evidence is convincing.
- Reversing a paired difference: With before-minus-after differences, a positive mean supports lower times after. With after-minus-before differences, the direction reverses. State the definition of the difference so the conclusion can be checked.
- Ignoring group order: In a two-sample test, the sign of \(\bar{x}_1-\bar{x}_2\) depends on which group is first. Keep the parameter order, alternative, calculation, and conclusion consistent.
- Calling failure to reject proof of no effect or equality: A large p-value means the data do not provide convincing evidence for the alternative at the chosen \(\alpha\). It does not establish that the null parameter value is true.
- Making the conclusion broader than the study: A conclusion should name the population the sampling method represents. A comparison in an observational study also does not, by itself, establish causation.
- Using vague language: Replace “there is a difference” with the specific direction, response, population, and units when the alternative is directional. A full-credit statement says whether the data provide convincing evidence for the population claim in context.
As a final check, read your conclusion alongside the hypotheses and the original claim. The parameter should describe the same quantity, the direction should match \(H_a\), and the evidence wording should follow the p-value decision. If any of those pieces disagree, revise the sentence rather than relying on a calculator’s label or a vague phrase such as “the claim is true.”
Check Your Understanding
For each situation, focus on whether the conclusion answers the stated claim precisely.
- A random sample of 20 lamps has a mean runtime above 10 hours. A right-tailed t test gives \(p=0.12\) at \(\alpha=0.05\). Write a conclusion about the claim that the population mean runtime exceeds 10 hours.
- For paired measurements, differences are defined as after minus before. What sign of the mean difference would support the claim that mean time is lower after an intervention?
- A two-sample test defines the parameter as \(\mu_A-\mu_B\), and the alternative is \(\mu_A-\mu_B>0\). In words, what population claim does this alternative represent?
- A test fails to reject \(H_0:\mu=25\). Explain why “the true mean is 25” is not a justified conclusion.
- An observational comparison finds convincing evidence of a lower mean response in one group. What additional design feature would be needed to make a causal conclusion more defensible?