From Computer Output to a Statistical Conclusion
A t procedure can produce several numbers at once. The challenge is to identify which number answers which question, and then explain what the output says about the population mean or means. As in “Writing the Full Two-Sample t Test Solution,” the calculation is only one part of a complete response. You must connect the output to the hypotheses, the confidence interval, and the study’s setting.
Software packages and calculators may arrange results differently, but the key quantities have the same meanings. The test statistic \(t\) measures how far the observed statistic is from the null value in estimated standard-error units. The degrees of freedom, \(df\), identify the t distribution used for the test and interval. The p-value measures how unusual a result at least this extreme would be if the null hypothesis were true. A confidence interval gives plausible values for the population parameter under the procedure’s conditions.
The displayed values may be rounded. A reported \(t=2.40\), for example, may have been calculated using more digits than shown. Use the output’s p-value and interval rather than trying to reproduce them from rounded values. A quick calculation with the sample summaries is still useful for detecting data-entry errors or a mismatch in group order.
A Reliable Reading Routine
Start by checking what procedure was run: a one-sample t test, a paired t test on differences, or an unpooled two-sample t test. Then locate the order of the data or groups. For two means, reversing the group order reverses the sign of \(t\) and the direction of the interval, even though the two-sided p-value stays the same.
Next, match each field to the question. A test output usually reports \(t\), \(df\), and \(p\). An interval output usually reports lower and upper endpoints, and may also report the sample estimate and margin of error. Some output gives a confidence interval alongside the test result; other software requires you to run a separate interval procedure.
Determine whether the output concerns one population mean, a mean of paired differences, or a difference between two population means.
Record the signed \(t\) statistic, \(df\), and p-value. Check that the alternative hypothesis and tail direction match the research question.
Record both endpoints and the confidence level. State the parameter and units represented by the interval, and check whether the null value is inside it when the test is two-sided.
Use the p-value and chosen \(\alpha\) to make the test decision, then describe the evidence about the population parameter. Keep the conclusion within the population or study units the design supports.
For a one-sample test, the statistic is based on \(\bar{x}-\mu_0\). For a paired test, it is based on the sample mean of the defined pairwise differences. For a two-sample test, the sign follows the order of the group means. These forms are useful checks, but the output is not a substitute for defining the parameter and checking the conditions, as in the earlier tutorials on t procedures.
Worked Example: Reading One-Sample t Output
Worked Example: Reading One-Sample t Output
A fictional manufacturer randomly selects 16 air filters from a production run to measure the number of months each lasts under a specified test. The sample mean is 18.4 months and the sample standard deviation is 4.0 months. Assume the production run contains at least 160 filters, and a plot of the lifetimes shows no severe skewness or extreme outliers. The manufacturer wants to test whether the population mean lifetime differs from 16 months at \(\alpha=0.05\). Software reports \(t=2.400\), \(df=15\), two-sided \(p=0.030\), and a 95% confidence interval of \((16.269,\ 20.531)\) months.
Let \(\mu\) be the true mean lifetime, in months, of filters from this production run under the specified test. The hypotheses are \(H_0:\mu=16\) months and \(H_a:\mu\ne16\) months.
Use a one-sample t test. The filters were randomly selected. The sample of 16 is no more than 10% of the run, since there are at least 160 filters. The plot shows no severe skewness or extreme outliers, supporting the t procedure for this sample size.
The output’s \(t=2.400\) means the sample mean is 2.400 estimated standard errors above the null mean. Check the statistic using the summaries:
The data provide convincing evidence that the true mean lifetime of filters from this production run under the specified test differs from 16 months. The interval gives plausible values for that mean; because the whole interval is above 16 months, it also indicates the direction of the difference.
The test and interval agree: the p-value is below 0.05, and 16 months is outside the matching 95% interval. As explained in “Connecting Confidence Intervals to Test Decisions,” this agreement is expected for a two-sided test at the 0.05 level and a 95% confidence interval for the same parameter. The interval does not say that 95% of individual filters last between 16.269 and 20.531 months.
Worked Example: Reading Output for Paired Data
Worked Example: Reading Output for Paired Data
A fictional physical therapy clinic records a flexibility score for each of 12 clients before and after a four-week stretching program. Define each difference as before score minus after score, so a positive difference means the score decreased after the program. The sample mean difference is 3.2 points, and the standard deviation of the differences is 4.0 points. The pairs were collected from a random sample of eligible clients; assume the clinic has at least 120 eligible clients. A plot of the differences shows no severe skewness or extreme outliers. Software reports \(t=2.771\), \(df=11\), two-sided \(p=0.018\), and a 95% confidence interval for the mean difference of \((0.659,\ 5.741)\) points.
This is paired output: the analysis is on the 12 differences, not on two independent samples of before and after scores. Let \(\mu_d\) be the true mean before-minus-after flexibility-score difference for eligible clients at this clinic. For a two-sided test of \(H_0:\mu_d=0\) against \(H_a:\mu_d\ne0\), the reported statistic can be checked:
The reported \(df=11\) matches \(n-1=12-1\) for a one-sample t procedure on the differences. The p-value is below 0.05, so reject \(H_0\). A suitable conclusion is: “The data provide convincing evidence that the true mean before-minus-after flexibility-score difference for eligible clients at this clinic is not zero.” Since the interval is entirely positive, the results point to a lower average score after the program, according to the stated difference order.
Do not describe the interval as an estimate of the mean before score or the mean after score separately. It estimates the population mean of the defined differences. If the subtraction order had instead been after minus before, \(t\) would be negative and the interval endpoints would change signs and reverse order, but the two-sided p-value and evidence of a difference would be unchanged.
Worked Example: Reading Output for Two Independent Means
Worked Example: Reading Output for Two Independent Means
A fictional transportation study randomly samples 16 commuters who usually travel by e-bike and 16 who usually travel by bus, then records commute time in minutes. The e-bike group has \(\bar{x}_1=42\) minutes and \(s_1=8\) minutes; the bus group has \(\bar{x}_2=38\) minutes and \(s_2=8\) minutes. Assume each population contains at least 160 commuters, the groups are independent, and plots show no severe skewness or extreme outliers in either group. The researchers test for a difference between the mean commute times. Unpooled software output, with e-bike listed first, reports \(t=1.414\), \(df=30\), two-sided \(p=0.168\), and a 95% confidence interval for \(\mu_1-\mu_2\) of \((-1.776,\ 9.776)\) minutes.
Let \(\mu_1\) and \(\mu_2\) be the true mean commute times for the e-bike and bus populations, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The random samples support inference to their respective populations, each sample is no more than 10% of its population, and the groups are independent. The plots support using the two-sample t procedure.
Check the test statistic using the group order shown:
The output’s \(df=30\) is the unpooled Welch degrees of freedom for these summaries. Since \(0.168>0.05\), fail to reject \(H_0\). The data do not provide convincing evidence of a difference in mean commute time between the two commuter populations. The interval includes zero, which represents no difference, and ranges from a possible e-bike mean that is 1.776 minutes lower to one that is 9.776 minutes higher than the bus mean.
A failure to reject is not proof that the means are equal. The interval shows that differences in either direction remain plausible under the procedure. Also, because the commuters were sampled from existing travel groups rather than randomly assigned to a travel method, a difference would not by itself establish that the method of travel caused a difference.
Common Mistakes and AP Exam Tips
- Reading a p-value as the probability the null hypothesis is true: The p-value is calculated assuming \(H_0\) is true. It is not the probability that \(H_0\) or \(H_a\) is correct.
- Ignoring the order of groups or differences: The sign of \(t\) and the interval’s direction depend on the order. State which group is first or how each paired difference was calculated.
- Reporting numbers without a parameter: A confidence interval is not just two endpoints. Say whether it estimates \(\mu\), \(\mu_d\), or \(\mu_1-\mu_2\), and include the context and units.
- Calling a non-significant result proof of no difference: When \(p>\alpha\), fail to reject \(H_0\). Say that the data do not provide convincing evidence for the alternative; do not say the means are equal.
- Assuming output checks the conditions: Software calculates from the entered data. It cannot establish that sampling was random, groups are independent, or plots support a t procedure. Check the design and data yourself.
- Mixing test and interval levels: For a two-sided test at level \(\alpha\), compare with the matching \(100(1-\alpha)\%\) interval for the same parameter. A different confidence level need not give the same decision.
For a full-credit AP response, use the output as evidence, not as the entire answer. Identify the procedure and parameter, report the relevant \(t\), \(df\), p-value, and interval accurately, make the reject-or-fail-to-reject decision using the stated \(\alpha\), and conclude in context. If rounded output seems slightly inconsistent, explain the result using the displayed values and avoid claiming exact agreement beyond their precision.
Check Your Understanding
Use the output fields, procedure, and context to answer each question.
- A one-sample t test reports \(t=-2.1\), \(df=19\), and two-sided \(p=0.049\). In words, what does the negative sign on \(t\) tell you, and what does the p-value mean?
- A paired analysis defines differences as after minus before and reports a 95% confidence interval of \((1.2,\ 4.8)\) points. What population parameter is being estimated, and what does the positive interval indicate?
- A two-sample output lists group B first and reports \(t=2.6\). If the group order is reversed, what happens to the sign of \(t\), and what happens to a two-sided p-value?
- A two-sided test at \(\alpha=0.05\) has \(p=0.12\), while a matching 95% confidence interval includes zero. State the decision and a correct conclusion about the evidence.
- Why is it not enough to copy \(t\), \(df\), and \(p\) from a computer output when writing a complete inference response?