Why a t Curve Needs Degrees of Freedom
In “The One-Sample t Test Statistic,” you learned to standardize the difference between a sample mean and a null value using the estimated standard error \(s/\sqrt{n}\). To judge how unusual that statistic would be if the null hypothesis were true, we use a t distribution. The particular t curve depends on its degrees of freedom, abbreviated \(df\).
For a one-sample t test or t interval about a population mean, the degrees of freedom are one less than the sample size. This rule applies to the one-sample procedures in this tutorial: if there are \(n\) observations, use \(df=n-1\). Degrees of freedom are not the number of observations, nor are they the observed t statistic. They identify which t distribution to use.
Why one less? The sample standard deviation \(s\) is calculated from the observations’ deviations from their sample mean \(\bar{x}\). Those deviations must add to zero. Once \(n-1\) of them are known, the last one is determined by that zero-sum requirement; it cannot vary freely. Thus, estimating \(\bar{x}\) uses one degree of freedom, leaving \(n-1\) degrees of freedom to estimate variability around the sample mean.
This is also why the sample variance uses \(n-1\) in its denominator when calculating \(s\). Under the Normal model used to derive the t procedure, standardizing by the estimated standard deviation produces a t statistic with \(n-1\) degrees of freedom. You do not need to derive the t distribution to use it: calculate \(n-1\), then use that \(df\) to select or draw the appropriate curve.
What the Degrees of Freedom Do to the Curve
A t curve is symmetric and centered at 0, just like the standard Normal curve. It has a bell shape, but its tails are heavier: compared with the standard Normal curve, more area lies farther from 0. The degrees of freedom determine how heavy those tails are. Smaller \(df\) means heavier tails and a wider curve; as \(df\) increases, the t curve gets closer to the standard Normal curve.
For a one-sample t procedure, \(n\) determines \(df\), and \(df\) determines which t curve to use. For example, a sample of 9 observations gives \(df=8\), while a sample of 25 gives \(df=24\). Both curves are centered at 0 and symmetric, but the curve with \(df=8\) has heavier tails.
The t statistic itself determines where the observed result sits on the horizontal axis. The alternative hypothesis determines which side or sides count as results at least as extreme as the observed statistic. A sketch helps make that distinction visible. The shaded area represents the p-value region; finding its numerical area using a calculator is the subject of the next tutorial.
Find the Degrees of Freedom and Shade the Test Curve
Identify the sample size \(n\) for the one-sample mean procedure.
Subtract one: \(df=n-1\). Use this value to identify the t curve.
Place it on the horizontal axis of the curve. The sign shows whether the sample mean is above or below the null value.
Shade the result or results at least as extreme as the observed statistic in the direction or directions specified by the alternative.
For a lower-tailed alternative \(H_a:\mu<\mu_0\), shade the area to the left of the observed t statistic. For an upper-tailed alternative \(H_a:\mu>\mu_0\), shade the area to its right. For a two-sided alternative \(H_a:\mu\ne\mu_0\), shade both tails beyond \(-|t|\) and \(|t|\). This two-tail rule works whether the observed statistic is positive or negative.
Worked Examples
Worked Example: An Upper-Tailed Test
A fictional orchard randomly selects 16 crates from a shipment of 400 crates to measure the mass of the fruit in each crate. The sample mean is 52 kilograms, and the sample standard deviation is 4 kilograms. The orchard tests whether the true mean crate mass is greater than 50 kilograms. Find the degrees of freedom, calculate the t statistic, and describe the shading.
State. Let \(\mu\) be the true mean fruit mass, in kilograms, per crate in this shipment. The hypotheses are \(H_0:\mu=50\) kilograms and \(H_a:\mu>50\) kilograms.
Plan and check conditions. The 16 crates were randomly selected, supporting the Random condition. Since selection was without replacement, check the 10% condition: \(0.10(400)=40\), and \(16\leq40\), so the sample is no more than 10% of the shipment and independence is reasonable. Because \(n=16<30\), the large-sample route is not met. The problem states that crate masses in this shipment follow an approximately Normal distribution, supporting the Normal/Large Sample condition for this small sample.
Do. Calculate the degrees of freedom and estimated standard error:
Then calculate the t statistic:
Conclude. Sketch a t curve centered at 0 and label it \(df=15\). Mark \(t=2.00\) to the right of 0 and shade the area to its right, because the alternative is \(H_a:\mu>50\). That shaded region is the p-value region. The statistic and sketch alone do not give its numerical area or establish whether there is convincing evidence that the mean crate mass exceeds 50 kilograms.
Worked Example: A Lower-Tailed Test
A fictional community kitchen randomly selects 9 containers from a delivery of 240 containers and measures the amount of soup in each. The sample mean is 7.5 liters, with a sample standard deviation of 3 liters. The kitchen tests whether the true mean amount is less than 8.5 liters. Find \(df\), calculate the statistic, and describe the curve.
State. Let \(\mu\) be the true mean amount of soup, in liters, in containers from this delivery. The hypotheses are \(H_0:\mu=8.5\) liters and \(H_a:\mu<8.5\) liters.
Plan and check conditions. The containers were randomly selected, supporting the Random condition. The 10% condition is met because \(0.10(240)=24\) and \(9\leq24\); independence is reasonable for this sample without replacement. Since \(n=9<30\), use the small-sample route for the Normal/Large Sample condition. The delivery process is described as approximately Normal, so this condition is supported.
Do. There are \(n=9\) observations, so
The estimated standard error is \(3/\sqrt{9}=1\) liter. Therefore,
Conclude. Draw a symmetric t curve centered at 0 and label it \(df=8\). Mark the observed statistic at \(t=-1.00\), to the left of 0, and shade the area to the left of \(-1.00\) because the alternative is lower-tailed. The smaller degrees of freedom give this curve heavier tails than the \(df=15\) curve in the previous example.
Worked Example: A Two-Sided Test
A fictional packaging facility randomly selects 25 sealed boxes from a day's production of 1,200 boxes. Their mean mass is 104 grams, and their sample standard deviation is 10 grams. The facility tests whether the true mean mass differs from 100 grams. Find \(df\), calculate the t statistic, and describe both shaded tails.
State. Let \(\mu\) be the true mean mass, in grams, of boxes produced that day. The hypotheses are \(H_0:\mu=100\) grams and \(H_a:\mu\ne100\) grams.
Plan and check conditions. The boxes were randomly selected, supporting the Random condition. The 10% condition is met because \(0.10(1200)=120\) and \(25\leq120\); independence is reasonable. Here \(n=25<30\), so the large-sample route is not met. The production process is described as approximately Normal, and the sample plot shows no pronounced outliers; this supports the Normal/Large Sample condition.
Do. The degrees of freedom are
The estimated standard error is
Thus, the observed test statistic is
Conclude. Sketch a t curve centered at 0 and label it \(df=24\). Mark \(t=2.00\) and shade both tails beyond \(-|t|=-2.00\) and \(|t|=2.00\). The alternative allows a mean either below or above 100 grams, so results at least as far from 0 in either direction count toward the p-value.
Degrees of Freedom for a Confidence Interval
A one-sample t confidence interval also uses \(df=n-1\). Instead of shading the p-value region for an observed statistic, a sketch for a two-sided interval shows the central confidence area between \(-t^*\) and \(t^*\). The leftover area is split equally between the two tails. The value \(t^*\) is the positive critical value chosen for the confidence level and degrees of freedom.
Worked Example: Sketching a 90% t Interval
A fictional library randomly selects 12 books from a collection of 900 books and measures their page counts. The sample mean is 72 pages and the sample standard deviation is 6 pages. Assume page counts are approximately Normally distributed, and the sample display shows no pronounced outliers. Identify the degrees of freedom and describe the area on a 90% t-interval sketch.
Plan and check conditions. The books were randomly selected, supporting the Random condition. The 10% condition is met because \(0.10(900)=90\) and \(12\leq90\); independence is reasonable. Since \(n=12<30\), the large-sample route is not met. The approximately Normal population description and the absence of pronounced outliers support the Normal/Large Sample condition.
Do. For this one-sample interval,
A 90% confidence level leaves \(1-0.90=0.10\) of the curve outside the central area. Dividing the remaining area equally gives \(0.10/2=0.05\) in each tail. With \(df=11\), the two-sided 90% critical value is approximately \(t^*=1.796\). A sketch therefore has area 0.05 to the left of \(-1.796\), area 0.90 between \(-1.796\) and \(1.796\), and area 0.05 to the right of \(1.796\).
The interval calculation uses the same \(df\) to select the critical value:
Conclude. The curve’s central 90% area corresponds to the 90% confidence level. In context, we are 90% confident that the true mean page count for books in this collection is between 68.89 and 75.11 pages. The calculation uses \(df=11\), not \(df=12\).
Common Mistakes and AP Exam Tips
- Using \(n\) degrees of freedom. For a one-sample t test or interval about a mean, calculate \(n-1\). If there are 16 observations, the procedure uses \(df=15\).
- Confusing degrees of freedom with the t statistic. The degrees of freedom select the curve; the t statistic locates the observed result on that curve. Report both when asked.
- Shading the wrong side. Follow the alternative hypothesis, not the sign of the statistic alone. A lower-tailed test shades left; an upper-tailed test shades right; a two-sided test shades both tails.
- Using the observed statistic as both two-sided boundaries. The two-sided boundaries are \(-|t|\) and \(|t|\). For an observed \(t=2.00\), they are \(-2.00\) and \(2.00\).
- Putting the confidence level in each tail. For a 90% two-sided interval, the central area is 0.90 and each tail has area \((1-0.90)/2=0.05\).
- Drawing every t curve the same way. All t curves are symmetric around 0, but curves with smaller degrees of freedom have heavier tails. Label the curve with the calculated \(df\).
For a full-credit curve sketch, label the horizontal axis \(t\), mark 0 at the center, write the correct degrees of freedom, and shade exactly the area called for by the alternative or confidence level. In a written explanation, state how \(df\) was calculated and connect the shaded region to the question being asked.
Check Your Understanding
Use \(df=n-1\) for each one-sample t procedure. For shading questions, describe the location and direction of the shaded area.
- A one-sample t test uses \(n=18\) observations. What degrees of freedom does it use?
- For \(H_a:\mu<\mu_0\), the observed statistic is \(t=-1.7\). Which part of the curve should be shaded?
- For \(H_a:\mu\ne\mu_0\), the observed statistic is \(t=-2.1\). State both tail boundaries for the shaded region.
- For a 90% two-sided t interval, what proportion of the curve lies in each tail?
- Which t curve has heavier tails: one with \(df=7\) or one with \(df=20\)?