From a t-Test Result to a Sentence in Context
In “What a P-Value Says About Sample Means,” you learned that a t-test p-value measures how unusual a test statistic would be if the null hypothesis were true. The next step is to express that idea in the setting of the problem. A strong interpretation names the population quantity in the null hypothesis, describes the observed sample result, and explains which results count as at least as extreme.
For a test about a single population mean, the sample result is \(\bar{x}\). For a test comparing two independent means, it is \(\bar{x}_1-\bar{x}_2\). For a paired t test, it is the sample mean of the pairwise differences, \(\bar{x}_d\). In each case, the t statistic measures how far the observed result is from the value specified by the null hypothesis, in estimated standard errors.
That sentence is a framework, not a script to recite without context. Replace general terms with the study’s population, variable, units, null value, observed sample result, and relevant direction. The p-value is calculated using the t distribution and the degrees of freedom for the procedure. It is not the probability that the null hypothesis is true.
When describing the sample result, keep the distinction between a statistic and a parameter clear. The sample mean is observed; the population mean is unknown. A test asks whether the observed sample result is surprising under a particular claim about that population mean. The p-value describes that surprise under the null model, not the chance that the sample mean itself is a population value.
One-Sample t Test: A Mean Above a Target
For a one-sided one-sample t test, translate the alternative into the direction of the sample results counted. If \(H_a:\mu>\mu_0\), results with t statistics at least as large as the observed statistic count. The contextual sentence should say that the sample mean is above the null value, while the probability refers to the test statistic under the null model.
Worked Example: Average Filling Time
A fictional packaging facility randomly selects 16 containers from a large production run and measures the time, in seconds, needed to fill each one. The sample mean is 54 seconds and the sample standard deviation is 8 seconds. The facility wants to know whether the true mean filling time is greater than 50 seconds. Assume the sample is less than 10% of the production run, and a plot of the data shows no strong skewness or outliers. Use \(\alpha=0.05\).
Let \(\mu\) be the true mean filling time for all containers in this production run. The hypotheses are \(H_0:\mu=50\) seconds and \(H_a:\mu>50\) seconds.
Use a one-sample t test. The containers were randomly selected, supporting inference to the production run. The 10% condition holds, supporting independence when sampling without replacement. Because \(n=16\), the data’s shape matters; the plot shows no strong skewness or outliers, so the t procedure is reasonable.
First find the standard error of the sample mean, then calculate the t statistic and right-tail probability. Since the alternative is greater than 50 seconds, count statistics at least as large as the observed one.
The degrees of freedom are \(16-1=15\). Using the \(t\) distribution with 15 degrees of freedom, the right-tail probability is \(P(T_{15}\geq 2.00)\approx0.0320\), rounded.
Because \(0.0320<0.05\), reject \(H_0\). The sample provides convincing evidence that the true mean filling time for containers in this production run is greater than 50 seconds.
A contextual interpretation of the p-value is: If the true mean filling time is 50 seconds, the probability of obtaining a t statistic of 2.00 or greater, in the direction of a higher mean, is about 0.0320. The observed sample mean is 54 seconds, which is above the null value. The p-value quantifies how often a result at least this far in the specified direction, relative to its estimated standard error, would occur under the null model.
Notice that the p-value sentence does not say “there is a 3.2% chance that the true mean is 50 seconds.” The calculation assumes that null value; it does not assign a probability to the value of \(\mu\).
Two Independent Means: State Both Groups
For an unpooled two-sample t test, the observed result is a difference in sample means. The contextual p-value sentence must preserve the group order: the statistic’s sign and the alternative hypothesis refer to group 1 minus group 2. As in “Identifying Two Independent Samples” and “Conditions for a Two-Sample t Test,” first use the study design and data to justify treating the groups as independent and using a t procedure.
Worked Example: Waiting Time at Two Service Desks
A fictional community center randomly samples eight visitors served at Desk A and eight different visitors served at Desk B. Waiting time is measured in minutes. At Desk A, the sample mean is 6 minutes and the sample standard deviation is 2 minutes. At Desk B, the sample mean is 8 minutes and the sample standard deviation is 2 minutes. Each sample is less than 10% of its visitor population, and plots for both groups show no strong skewness or outliers. The question is whether the true mean wait at Desk A is shorter.
Let \(\mu_A\) be the true mean waiting time for visitors served at Desk A, and let \(\mu_B\) be the true mean waiting time for visitors served at Desk B. The hypotheses are \(H_0:\mu_A-\mu_B=0\) and \(H_a:\mu_A-\mu_B<0\). The alternative is left-tailed because a shorter mean at Desk A means \(\mu_A-\mu_B\) is negative.
Use an unpooled two-sample t test. The visitors were randomly sampled, and the two groups contain different visitors, so the samples are independent rather than paired. The 10% condition holds for each sample. With eight visitors in each group, the plots are relevant to the Nearly Normal condition; the stated lack of strong skewness or outliers supports using the procedure.
The observed difference is \(6-8=-2\) minutes. As in “The Two-Sample t Test Statistic,” calculate the standard error using each group’s variance contribution separately:
The Welch degrees of freedom are 14. For the left-tailed alternative, the p-value is \(P(T_{14}\leq-2.00)\approx0.0326\), rounded. In context: If the true mean waiting times at the two desks are equal, the probability of obtaining a t statistic of \(-2.00\) or less, indicating a difference at least as far in the direction of a shorter mean wait at Desk A, is about 0.0326.
Because \(0.0326<0.05\), reject \(H_0\). The data provide convincing evidence that the true mean waiting time at Desk A is shorter than the true mean waiting time at Desk B. This conclusion is about population means, not just the sample means of 6 and 8 minutes.
Paired t Test: Interpret the Mean Difference
A paired t test is a one-sample t test applied to the differences within pairs. Define the difference in a clear order before writing hypotheses or interpreting the p-value. For example, if \(d=\text{after}-\text{before}\), then a negative mean difference indicates a decrease. The p-value sentence should describe the population mean of those differences and the direction specified by the alternative.
Worked Example: Time to Complete a Practice Route
A fictional recreation program randomly selects nine participants and records each person’s time to complete a practice route before and after a training session. Define each difference as \(d=\text{after time}-\text{before time}\), in minutes. The sample mean difference is \(-2.0\) minutes and the sample standard deviation of the differences is 3.0 minutes. Assume the participants are less than 10% of the target population and a plot of the differences shows no strong skewness or outliers. The program asks whether the population mean change is negative.
Let \(\mu_d\) be the true mean after-minus-before change in route completion time for participants in the target population. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\). The same participants were measured twice, so the observations are paired; the analysis uses the nine differences, not two independent-sample means.
Use a one-sample t test on the differences. The participants were randomly selected, supporting inference to the target population. The 10% condition supports independence among sampled participants, and the plot of the differences supports the Nearly Normal condition for this small sample.
The degrees of freedom are \(9-1=8\). Because the alternative is left-tailed, the p-value is \(P(T_8\leq-2.00)\approx0.0403\), rounded. In context: If the true mean after-minus-before change in route completion time is zero, the probability of obtaining a t statistic of \(-2.00\) or less, in the direction of a decrease, is about 0.0403.
At \(\alpha=0.05\), reject \(H_0\). The data provide convincing evidence that the true mean change in route completion time for the target population is negative. Because this example describes measurements before and after training, the test result alone does not establish that training caused the change; the study design also matters.
Two-Sided Tests Count Both Directions
A two-sided alternative asks whether a population mean differs from a null value, without specifying which direction in advance. Its p-value counts results at least as far from the null value as the observed result in either direction. The contextual sentence must therefore mention both directions, even if the sample mean lies on just one side of the null value.
Worked Example: Average Time to Complete a Form
A fictional software team randomly selects nine users from a large pool and records how many minutes each takes to complete a form. The sample mean is 32 minutes and the sample standard deviation is 6 minutes. The team wants to test whether the true mean completion time differs from 30 minutes. Assume the sample is less than 10% of the pool and the data show no strong skewness or outliers.
Let \(\mu\) be the true mean completion time for users in the pool. The hypotheses are \(H_0:\mu=30\) minutes and \(H_a:\mu\ne30\) minutes. Use a one-sample t test: the users were randomly sampled, the 10% condition holds, and the data’s shape supports the t procedure for this small sample.
For the two-sided alternative, both tails count: \(P(T_8\leq-1.00)+P(T_8\geq1.00)\approx0.3466\), rounded. In context: If the true mean completion time is 30 minutes, the probability of obtaining a t statistic of \(-1.00\) or less, or \(1.00\) or greater, is about 0.3466. The sample mean is 32 minutes, but results in the opposite direction also count because the question asks about a difference in either direction.
Because \(0.3466>0.05\), fail to reject \(H_0\). These data do not provide convincing evidence that the true mean completion time differs from 30 minutes. This is not proof that the population mean equals 30 minutes; it means this test does not provide convincing evidence against that null value.
Common Mistakes and AP Exam Tips
- Leaving out the null assumption: A p-value is calculated assuming \(H_0\) is true. Include the null value or null relationship in the sentence.
- Interpreting the p-value as a probability about the parameter: Do not write “There is a 4% chance that the means are equal.” Instead, describe the probability of a test statistic at least as extreme as the observed one under the null model.
- Reporting only “as extreme” without its direction: For a one-sided test, state whether the relevant statistics are larger or smaller. For a two-sided test, say that results in either direction count.
- Using the sample mean where the population mean belongs: Sample means such as \(\bar{x}\) are observed statistics. Define the population parameter in context, then use the sample result to explain what was observed.
- Ignoring the order of a difference: In a two-sample or paired test, preserve the group order or difference definition used in the hypotheses. Reversing it changes the sign and can change which tail is relevant.
- Confusing the p-value with the test decision: State the p-value, compare it with \(\alpha\), and then conclude in context. A p-value is not itself the conclusion about convincing evidence.
- Treating a non-significant result as proof of no difference: “Fail to reject” means the data do not provide convincing evidence against the null at the chosen significance level; it does not establish equality.
As discussed in “Linking Two-Sample Intervals and Tests,” a two-sided test and a matching confidence interval can be used to assess the same null value. They answer related but distinct questions: the test evaluates evidence against a specific null value, while the interval gives a range of plausible values for the population parameter. Do not describe the p-value as a confidence interval or as the probability that a parameter lies in one.
Check Your Understanding
For each question, focus on the parameter in the null hypothesis and the results counted by the alternative.
- A one-sample t test has \(H_0:\mu=12\), \(H_a:\mu>12\), and \(t=1.7\). Which part of the t distribution is used for the p-value, and what null assumption belongs in the contextual sentence?
- A two-sided test comparing two population means has \(t=-2.1\). Describe which t statistics count as at least as extreme as the observed statistic.
- For a paired test, differences are defined as after minus before, and \(H_a:\mu_d<0\). What does a negative sample mean difference represent in context?
- Explain why “the p-value is the probability that the two population means are equal” is incorrect.
- A t test gives \(p=0.08\) with \(\alpha=0.05\). Write a conclusion that correctly distinguishes the test decision from a claim that the population means are equal.