From the Test Statistic to Calculator Output
In “The Two-Sample t Test Statistic” and “Reading Degrees of Freedom From 2-SampTTest Output,” you learned how the test statistic measures the difference between sample means in estimated standard-error units, and how the calculator’s unpooled procedure uses Welch degrees of freedom. This tutorial connects those ideas to the calculator’s menus and output.
A 2-SampTTest compares two population means for the same quantitative variable. You can enter the individual observations in lists, or enter each group’s sample mean, sample standard deviation, and sample size. In either mode, make sure Group 1 and Group 2 match the order in your hypotheses. The sign of \(t\), and the tail used to calculate the p-value, depend on that order.
Setting Up 2-SampTTest
On a TI-84, open the STAT menu, select TESTS, and choose 2-SampTTest. Menu wording can differ slightly on other calculators, but the required choices are the same.
Choose Data when you have entered the individual observations in two lists. Enter the list names, such as L1 and L2, and check that the values in each list belong to the intended group. Choose Stats when the problem gives summary statistics. Enter \(\bar{x}_1\), \(s_1\), and \(n_1\) for Group 1, and \(\bar{x}_2\), \(s_2\), and \(n_2\) for Group 2.
For the standard AP Statistics two-sample t test, set Pooled: No. This is the unpooled test, which estimates the two groups’ variance contributions separately and uses Welch degrees of freedom. Do not select a pooled test unless a task specifically calls for it.
Choose the alternative to match the research question. Use \(\mu_1\ne\mu_2\) for a difference in either direction, \(\mu_1<\mu_2\) for a lower mean in Group 1, or \(\mu_1>\mu_2\) for a higher mean in Group 1. The null hypothesis is \(H_0:\mu_1-\mu_2=0\). Define both population means in context, as in “Defining Both Population Means in Context.”
The calculator’s output commonly includes \(\bar{x}_1\), \(\bar{x}_2\), \(Sx_1\), \(Sx_2\), \(n_1\), \(n_2\), \(t\), \(p\), and \(df\). Here, \(Sx_1\) and \(Sx_2\) are the sample standard deviations. The group means and sample sizes are useful checks that the data were entered correctly. The test statistic and degrees of freedom should agree with the formulas covered earlier in this unit; the p-value is calculated from the selected alternative and the reported t distribution.
A Four-Step Calculator Workflow
Define \(\mu_1\) and \(\mu_2\) in context, and write \(H_0:\mu_1-\mu_2=0\) with the alternative that answers the research question.
Check that the response is quantitative, the samples are independent, and the design supports inference. Check the 10% condition for each sample drawn without replacement, and assess whether the data support a t procedure.
Enter the lists or summary statistics, select the correct alternative, and choose Pooled: No. Read and report the group means, t statistic, p-value, and Welch degrees of freedom.
Compare the p-value with the significance level, state whether you reject or fail to reject \(H_0\), and explain what the evidence says about the population means in context.
The calculator does not check the study design or data shape for you. As in “Conditions for a Two-Sample t Interval,” consider random selection or random assignment, independence within and between groups, the 10% condition when relevant, and whether the samples are large enough or the data show no strong skewness or outliers. For small samples, examine the distributions before using a t procedure.
Worked Examples
Worked Example: Entering Individual Data in Lists
A hypothetical lab compares the time, in minutes, needed to complete a short task using two different interfaces. Four independent randomly selected users try each interface. The recorded times are:
| Interface 1 | Interface 2 |
|---|---|
| 4, 6, 8, 10 | 3, 5, 9, 11 |
State: Let \(\mu_1\) be the true mean task-completion time for users of Interface 1 and \(\mu_2\) the true mean for users of Interface 2. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2\ne0\), since the question asks whether the mean times differ.
Plan: The response is quantitative, and different users use the two interfaces, so the samples are independent rather than paired. Suppose the users were randomly selected from large user populations, and each sample is less than 10% of its population. Both samples are small, so inspect the values: each list is symmetric around its mean and has no apparent outlier. These checks support using an unpooled two-sample t test for this example.
Do: Enter the Interface 1 values in L1 and the Interface 2 values in L2. Choose Data, set List1 to L1 and List2 to L2, select \(\mu_1\ne\mu_2\), and set Pooled: No. The calculator reports \(\bar{x}_1=7\), \(\bar{x}_2=7\), \(n_1=n_2=4\), \(t=0\), \(df=5.4\), and \(p=1\).
Check the key output. Each group’s mean is 7 minutes. The sample variances are \(20/3\) and \(40/3\), so the standard error is:
The test statistic is \((7-7)/2.2361=0\). Welch’s degrees of freedom are:
With a two-sided alternative and \(t=0\), the p-value is 1: the statistic is exactly at the center of the null t distribution.
Conclude: At the 5% significance level, fail to reject \(H_0\). These samples provide no evidence of a difference in the population mean task-completion times. This very small example illustrates calculator entry and output; the conditions and study context still matter for any real conclusion.
Worked Example: Entering Summary Statistics and Reading the Test
A hypothetical transit comparison examines passenger waiting time, in minutes, at two types of stops. Independent random samples give these summary statistics: Stop Type 1 has \(n_1=8\), \(\bar{x}_1=5\), and \(s_1=2\); Stop Type 2 has \(n_2=12\), \(\bar{x}_2=7\), and \(s_2=3\). The question is whether the mean wait is shorter at Stop Type 1.
State: Let \(\mu_1\) and \(\mu_2\) be the true mean passenger waiting times at Stop Type 1 and Stop Type 2, respectively. Test \(H_0:\mu_1-\mu_2=0\) against \(H_a:\mu_1-\mu_2<0\).
Plan: Waiting time is quantitative. The groups consist of different stops and are not paired. Assume the stops were randomly selected from large populations, with each sample less than 10% of its population. Because both samples are small, suppose plots of the observations show no strong skewness or outliers. These conditions support an unpooled two-sample t test.
Do: Choose Stats mode. Enter \(5,2,8\) for \(\bar{x}_1,s_1,n_1\), and \(7,3,12\) for \(\bar{x}_2,s_2,n_2\). Select \(\mu_1<\mu_2\) and Pooled: No. The calculator reports \(t\approx-1.7889\), \(df\approx17.9907\), and a left-tail \(p\)-value of \(0.04524\), rounded. It also displays the group means, 5 and 7 minutes, and the sample sizes.
The standard error and statistic verify the calculator output:
For Welch degrees of freedom, the variance contributions are \(a=2^2/8=0.5\) and \(b=3^2/12=0.75\). Substitution gives:
The p-value is the probability, assuming \(H_0\) is true, of obtaining a t statistic at least as far in the direction of \(H_a\) as \(-1.7889\). Here that probability is about \(0.04524\).
Conclude: At the 5% significance level, reject \(H_0\), since \(0.04524<0.05\). The samples provide convincing evidence that the mean passenger waiting time is shorter at Stop Type 1 than at Stop Type 2. The evidence supports a difference in the stated direction; it does not establish how large the population difference is.
Worked Example: Keeping the Group Order and Tail Aligned
Use the transit summary statistics from the previous example, but enter Stop Type 2 first and Stop Type 1 second. Now Group 1 has mean 7 minutes and Group 2 has mean 5 minutes. The research question remains whether Stop Type 1 has the shorter mean wait.
State: Keep the parameter definitions tied to the calculator’s new order: \(\mu_1\) is the true mean for Stop Type 2 and \(\mu_2\) is the true mean for Stop Type 1. The same claim that Stop Type 1 has the shorter mean is now \(H_a:\mu_1-\mu_2>0\), with \(H_0:\mu_1-\mu_2=0\).
Plan: The study design and conditions are unchanged. Use Stats mode, put the Stop Type 2 summaries in Group 1 and Stop Type 1 summaries in Group 2, select \(\mu_1>\mu_2\), and keep Pooled: No.
Do: The difference in sample means is now \(7-5=2\) minutes, so the standard error remains \(\sqrt{1.25}\approx1.1180\) minutes and \(t\approx1.7889\). The Welch degrees of freedom are unchanged at about \(17.9907\), because the same sample sizes and standard deviations are used. With the upper-tail alternative \(\mu_1>\mu_2\), the calculator gives \(p\approx0.04524\).
If you instead select \(\mu_1<\mu_2\) with this reversed group order, the test uses the opposite tail and gives \(p\approx0.95476\), the complement of \(0.04524\). The group order and alternative must be changed together to test the same research claim.
Conclude: With the correctly aligned upper-tail alternative, the evidence and conclusion match the previous example: at the 5% level, reject \(H_0\) and conclude there is convincing evidence that Stop Type 1 has a shorter population mean wait than Stop Type 2. Reversing the group order changes the sign of \(t\), not the underlying comparison, when the hypotheses are also expressed consistently.
Common Mistakes and AP Exam Tips
- Entering data in the wrong group: Check the displayed \(\bar{x}_1\) and \(\bar{x}_2\) against the intended groups before interpreting \(t\). The sign of the statistic follows Group 1 minus Group 2.
- Choosing the wrong tail: The calculator does not know the research question. Select the alternative before reading \(p\), and confirm that the test direction matches the hypotheses.
- Using pooled variances by default: For the standard AP unpooled two-sample t test, choose Pooled: No.
- Calling the p-value the probability that the null hypothesis is true: A p-value is calculated assuming the null hypothesis is true. It is the probability of a statistic at least as extreme as the observed one in the direction specified by the alternative.
- Treating \(df\) as the number of observations: The calculator’s decimal \(df\) is the Welch degrees of freedom. Report it as shown; it is not \(n_1+n_2\).
- Reporting output without context: A full-credit conclusion identifies the two population means, the direction or nature of the evidence, and the significance level. Calculator output alone is not a conclusion.
- Assuming the calculator checked conditions: You must justify the design, independence, 10% condition when applicable, and suitability of the t procedure from the study information.
Check Your Understanding
Answer each question by connecting calculator setup and output to the two-mean test.
- When should you choose Data mode, and when should you choose Stats mode?
- A calculator reports \(t=-2.1\) after Group 1 was entered first. What does the negative sign say about the observed difference in sample means?
- For the standard unpooled two-sample t test, which pooled setting should you select?
- A test reports \(p=0.03\) at significance level \(0.05\). State the decision and what the p-value means under the null hypothesis.
- If you reverse the group order, what must you do to the alternative hypothesis to test the same research claim?