Tutorials › AP Statistics › Interpreting Calculator Output for a t Test

P-values and mean-inference conclusions · Tutorial 756 of 1000

Interpreting Calculator Output for a t Test

Practice identifying the test statistic, degrees of freedom, and p-value on calculator screens, then use them to make a correct conclusion in context.

Intermediate 10 min read

What You'll Learn

  • Identify the procedure and parameter represented by a calculator’s t-test output.
  • Interpret the displayed \(t\) statistic, degrees of freedom, and p-value.
  • Match the calculator’s alternative-hypothesis setting to the research question.
  • Check a displayed t statistic using sample summaries without rounding too early.
  • Write a reject-or-fail-to-reject conclusion in context for one-sample, paired, and two-sample tests.

Reading a t-Test Screen

A calculator can report a test statistic, degrees of freedom, and p-value in a compact screen. Reading those numbers correctly requires more than recognizing the labels: you need to know which parameter was tested, which alternative was selected, and what the study design allows you to conclude. This tutorial builds on “Conclusions From Computer Output for Means” by focusing on the values and settings commonly displayed by a calculator.

On a TI-84, a one-sample or paired t test typically reports the null mean \(\mu_0\), \(t\), \(p\), \(df\), \(\bar{x}\), \(S_x\), and \(n\). For paired data, \(\bar{x}\) and \(S_x\) refer to the defined differences. A two-sample t test typically reports the sample summaries for both groups, followed by \(t\), \(p\), \(df\), and the pooled setting. Other calculators may use different labels or layouts, but the quantities have the same roles.

Definition: In t-test output, \(t\) is the signed number of estimated standard errors between the observed statistic and the value specified by \(H_0\). The degrees of freedom, \(df\), select the t distribution used to calculate the p-value. The p-value is calculated under \(H_0\), using the alternative hypothesis and the observed \(t\).

The sign of \(t\) gives the direction of the observed result relative to the null value and the order of the data. It does not, by itself, tell you whether the result is statistically significant. The p-value depends on the alternative entered in the calculator: a left-tailed, right-tailed, or two-sided test uses a different tail area. As discussed in “One-Sided and Two-Sided P-Values Compared,” the alternative must match the research question.

A Routine for Interpreting Calculator Output

Before interpreting any field, identify the procedure. A one-sample t test concerns one population mean; a paired t test is a one-sample t test on defined pairwise differences; and an unpooled two-sample t test concerns the difference between two population means. The parameter and the data order determine what the sign of \(t\) means.

1
Identify the procedure and parameter.
Determine whether the output concerns \(\mu\), a mean difference \(\mu_d\), or a difference \(\mu_1-\mu_2\). For a two-sample test, note which group is listed first.
2
Check the alternative setting.
Confirm whether the calculator used \(<\), \(>\), or \(\ne\). A correct statistic paired with the wrong alternative produces a p-value for the wrong research question.
3
Read the test results.
Record \(t\), \(df\), and \(p\). Treat a displayed p-value as rounded, and use the calculator’s p-value rather than recalculating it from rounded output.
4
Make and explain the decision.
Compare \(p\) with the stated \(\alpha\), then reject or fail to reject \(H_0\). State what the data show about the population parameter in context.

The sample summaries on the screen provide useful checks. For a one-sample or paired test, the standard error is \(s/\sqrt{n}\), and the test statistic compares the sample mean with the null value. For an unpooled two-sample test, the standard error uses both sample variances and sample sizes. Keep full calculator precision in intermediate steps: rounding a standard error too early can make a reconstructed \(t\) differ slightly from the screen.

Key takeaway: A calculator reports results for the procedure and alternative you selected; it does not define the population parameter, check the study design, or write the conclusion. Interpret its values only after verifying those pieces.

Worked Example: One-Sample t-Test Output

Worked Example: One-Sample t-Test Output

A fictional greenhouse randomly selects 25 seedlings from a large production batch and records their heights, in centimeters, after a specified growing period. The sample mean is 52.4 cm and the sample standard deviation is 8.0 cm. Assume the batch contains at least 250 seedlings, and a plot of the heights shows no severe skewness or extreme outliers. The greenhouse tests whether the batch’s mean height exceeds 50 cm, using \(\alpha=0.05\). A calculator screen displays:

Screen fieldDisplayed value
\(\mu_0\)50
\(t\)1.500
\(p\)0.0733
\(df\)24
\(\bar{x}\), \(S_x\), \(n\)52.4, 8.0, 25
Alternative setting\(\mu>\mu_0\)

Let \(\mu\) be the true mean height, in centimeters, of seedlings in this production batch after the specified growing period. The null and alternative hypotheses are \(H_0:\mu=50\) cm and \(H_a:\mu>50\) cm. The output’s \(\mu_0=50\) is the null value, not the sample mean; \(\bar{x}=52.4\) is the sample mean.

Check the statistic using the displayed summaries. The standard error is \(8.0/\sqrt{25}=1.6\) cm, so

$$ t=\frac{\bar{x}-\mu_0}{s/\sqrt{n}} =\frac{52.4-50}{8.0/\sqrt{25}} =\frac{2.4}{1.6} =1.500. $$

The value \(t=1.500\) means the sample mean is 1.500 estimated standard errors above the null mean. The \(df=24\) matches \(n-1=25-1\). Because the selected alternative is right-tailed, the displayed \(p=0.0733\) is the area to the right of \(t=1.500\) under a t distribution with 24 degrees of freedom.

Since \(0.0733>0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence that the true mean height of seedlings in this batch exceeds 50 cm. This is not evidence that the mean equals 50 cm; it means the output does not show sufficiently strong evidence for the stated greater-than alternative at the chosen significance level.

Worked Example: Paired t-Test Output

Worked Example: Paired t-Test Output

A fictional community garden measures the time, in minutes, that each of 12 volunteers needs to complete a standard planting task, both before and after a brief tool-use workshop. Define each difference as before time minus after time, so a positive difference means the volunteer took longer before the workshop. The volunteers are a random sample of eligible garden volunteers; assume there are at least 120 eligible volunteers. A plot of the differences shows no severe skewness or extreme outliers. The garden tests whether the mean before-minus-after difference is below zero, using \(\alpha=0.05\). The calculator’s one-sample T-Test screen, run on the differences, displays \(\mu_0=0\), \(t=-2.598\), \(p=0.0124\), \(df=11\), \(\bar{x}=-3.0\), \(S_x=4.0\), and \(n=12\), with the alternative \(\mu_d<0\).

1
State.
Let \(\mu_d\) be the true mean before-minus-after task-time difference, in minutes, for eligible volunteers at this garden. The hypotheses are \(H_0:\mu_d=0\) and \(H_a:\mu_d<0\). A negative difference corresponds to a shorter time after the workshop.
2
Plan and check conditions.
Use a paired t test, which is a one-sample t test on the 12 differences. The volunteers were randomly sampled. The sample is no more than 10% of the eligible volunteers because there are at least 120 and \(12/120=0.10\). The plot of the differences shows no severe skewness or extreme outliers, supporting a t procedure for this sample size.
3
Do.
The screen’s \(\bar{x}=-3.0\) and \(S_x=4.0\) summarize the differences, not the before and after measurements separately. Check \(t\) without rounding the standard error first:
$$ SE_{\bar d}=\frac{s_d}{\sqrt{n}}=\frac{4.0}{\sqrt{12}}, \qquad t=\frac{\bar d-0}{SE_{\bar d}} =\frac{-3.0}{4.0/\sqrt{12}} =-\frac{3\sqrt{12}}{4} \approx-2.598. $$
The \(df=11\) matches \(n-1=12-1\). The selected alternative is left-tailed, and the calculator reports \(p=0.0124\). Since \(0.0124<0.05\), reject \(H_0\).
4
Conclude in context.
The data provide convincing evidence that the true mean before-minus-after task-time difference for eligible volunteers at this garden is below zero. In context, the results support a lower mean task-completion time after the workshop for this population.

The minus sign on \(t\) agrees with the negative sample mean difference and with the direction of \(H_a\). If the differences had instead been defined as after minus before, the sign of \(t\) would reverse and the alternative would need to reverse too. The paired structure is essential: the calculator analyzed one difference for each volunteer, not two independent groups of times.

Worked Example: Two-Sample t-Test Output

Worked Example: Two-Sample t-Test Output

In a fictional usability experiment, 25 participants are randomly assigned to each of two website layouts. The response is the time, in seconds, needed to find a specified setting. Layout A has \(\bar{x}_1=44\) seconds and \(s_1=10\) seconds; Layout B has \(\bar{x}_2=38\) seconds and \(s_2=10\) seconds. Assume the response plots show no severe skewness or extreme outliers. The experiment tests for a difference in mean times. With Layout A entered as group 1, an unpooled two-sample t test reports \(t=2.121\), \(df=48\), and two-sided \(p=0.0391\).

Let \(\mu_1\) and \(\mu_2\) be the true mean setting-finding times, in seconds, for participants assigned to Layout A and Layout B, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\). The randomized assignment supports comparing the treatment groups. The groups are independent because each participant was assigned to only one layout. Each group has 25 participants. Because the participants were randomly assigned, the 10% condition is not needed here. The plots show no severe skewness or extreme outliers, which supports using the two-sample t procedure.

Check the statistic using the group order shown on the screen:

$$ SE=\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}} =\sqrt{\frac{10^2}{25}+\frac{10^2}{25}} =\sqrt{8} \approx2.8284\text{ seconds}, \qquad t=\frac{44-38}{\sqrt{8}} \approx2.121. $$

Because the two sample sizes and sample standard deviations are equal, the unpooled Welch degrees of freedom are \(48\), matching \(25+25-2=48\) in this particular case. The calculator’s two-sided p-value is \(0.0391\). Since \(0.0391<0.05\), reject \(H_0\). The data provide convincing evidence of a difference in mean setting-finding time between participants assigned to the two layouts. The positive \(t\) indicates that the observed sample mean is higher for Layout A; the result does not mean the p-value is the probability that the null hypothesis is true.

For this randomized experiment, the evidence concerns the mean response associated with assignment to the two layouts. Be careful not to reverse the group order when describing the sign: if Layout B were entered first, the test statistic would be negative instead. The two-sided p-value would remain the same.

Common Mistakes and AP Exam Tips

  • Using the sign of \(t\) as the decision: A positive or negative statistic indicates direction relative to the null value and data order. Compare the p-value with \(\alpha\) to decide whether to reject.
  • Ignoring the alternative setting: A calculator may return a p-value for the wrong tail if the wrong alternative was selected. Check the setting against \(H_a\) before interpreting \(p\).
  • Misidentifying the parameter: In paired output, the screen’s \(\bar{x}\) is the mean of the defined differences. In two-sample output, the group order determines whether the estimate is \(\bar{x}_1-\bar{x}_2\) or its reverse.
  • Treating degrees of freedom as the sample size: For a one-sample or paired t test, \(df=n-1\). An unpooled two-sample test may report decimal Welch degrees of freedom, depending on the sample summaries.
  • Rounding too early when checking \(t\): Keep the full standard-error expression in the denominator until the final step. For example, calculate \(2/(2/\sqrt{10})=\sqrt{10}\), rather than dividing by a rounded standard error and expecting identical final digits.
  • Copying the screen without a contextual conclusion: Full-credit communication identifies the parameter, compares \(p\) with the stated \(\alpha\), makes the reject-or-fail-to-reject decision, and describes the evidence in context.
  • Claiming that failure to reject proves equality: If \(p>\alpha\), say the data do not provide convincing evidence for the alternative. Do not say that the population means are equal.

A useful final check is to make sure the reported \(t\), direction of the alternative, and conclusion tell a consistent story. A displayed p-value is rounded, so a small difference from a calculation using rounded summaries is not automatically an error. Explain the output using its reported precision, while making the decision from the stated \(\alpha\).

Key takeaway: Identify the test and parameter first, then read \(t\), \(df\), and \(p\) in light of the calculator’s alternative setting. Use the p-value and \(\alpha\) to make the decision, and finish with a conclusion about the population parameter in context.

Check Your Understanding

For each item, identify what the output means and how it should affect the conclusion.

  1. A one-sample test of \(H_a:\mu<18\) reports \(t=-1.9\), \(df=14\), and \(p=0.039\). What does the negative sign indicate, and what does the p-value represent?
  2. A paired calculator screen reports \(\bar{x}=2.5\) and \(S_x=3.0\). The differences were defined as after minus before. Which data do these summary statistics describe?
  3. A two-sample output lists Group A first and reports \(t=-2.2\). Which sample mean is lower, and what would happen to the sign if the group order were reversed?
  4. A test at \(\alpha=0.05\) reports \(p=0.08\). State the decision and write the key evidence conclusion without claiming the null hypothesis is true.
  5. Why should you check the alternative-hypothesis setting before interpreting a calculator’s p-value?