What Does a P-Value Measure?
A t test starts with a claim about one or more population means and asks whether the sample results are surprising if that claim is true. The p-value measures how surprising the test result is under the null hypothesis. To understand it, connect the observed sample mean—or difference in sample means—to the t statistic and the t distribution.
As in “Finding the P-Value for a Two-Sample t Test,” the alternative hypothesis determines which tail area to use. Here we focus on what that area says about sample means. A p-value does not tell us the probability that the null hypothesis is true.
The p-value is calculated from a probability model for results that could occur over repeated sampling if the null hypothesis were true. For a t test, that model is a t distribution. A test statistic tells us how far the observed sample result is from the value expected under the null, measured in estimated standard errors.
For a one-sample t test of \(H_0:\mu=\mu_0\), the observed result is \(\bar{x}\), and the estimated standard error is \(s/\sqrt{n}\). For an unpooled two-sample t test of \(H_0:\mu_1-\mu_2=0\), the observed result is \(\bar{x}_1-\bar{x}_2\), and the estimated standard error is \(\sqrt{s_1^2/n_1+s_2^2/n_2}\). In either case, the p-value comes from the t distribution for the procedure, using its degrees of freedom.
“As Extreme” Depends on the Alternative
A left-tailed alternative asks whether the parameter is less than the null value, so results at least as extreme as the observed statistic lie to its left. A right-tailed alternative asks whether it is greater, so the relevant area is to the right. A two-sided alternative looks for a difference in either direction; results at least as extreme are those whose distance from zero is at least as large as the observed statistic’s distance.
For example, if a two-sided test produces \(t=2\), statistics at least as extreme are \(t\leq-2\) or \(t\geq2\). If a right-tailed test produces \(t=2\), only results at or above 2 count. If the alternative points left but the observed statistic is positive, the p-value is the area to the left of that positive statistic; it may be large because the observed result points away from the alternative.
For a t distribution with \(df\) degrees of freedom, technology can find these tail areas using a command such as tcdf. For a two-sided test, the area is \(P(T\leq-|t_{\text{obs}}|)+P(T\geq|t_{\text{obs}}|)\). Because the t distribution is symmetric, this is also twice the area beyond \(|t_{\text{obs}}|\) in one tail.
Worked Example: A Sample Mean Above a Target
A fictional materials lab randomly selects 16 sheets from a large production run and measures their thickness in millimeters. The sample mean is 52.0 mm and the sample standard deviation is 8.0 mm. The lab wants to test whether the population mean thickness is greater than 50 mm, using \(\alpha=0.05\). Assume the sample is less than 10% of the production run and a plot shows no strong skewness or outliers.
Let \(\mu\) be the true mean thickness of all sheets in the production run. The hypotheses are \(H_0:\mu=50\) mm and \(H_a:\mu>50\) mm.
Use a one-sample t test. The sheets were randomly selected, and the 10% condition holds, supporting independence when sampling without replacement. Because \(n=16\), the plot is relevant: it shows no strong skewness or outliers, so using a t procedure is reasonable.
The standard error estimates the standard deviation of the sample mean under repeated sampling. The observed sample mean is 1.0 estimated standard error above the null value.
The degrees of freedom are \(16-1=15\). Since \(H_a:\mu>50\), the p-value is the area to the right of \(t=1.00\), not a two-sided area. A t calculation gives \(p=P(T_{15}\geq1.00)\approx0.1666\), rounded.
Because \(0.1666\) is greater than \(0.05\), fail to reject \(H_0\). The sample does not provide convincing evidence that the true mean sheet thickness in this production run is greater than 50 mm.
The p-value is not the chance that the true mean equals 50 mm. It is the probability, assuming a true mean of 50 mm and the t-test conditions, of obtaining a test statistic of at least 1.00 in the direction of the alternative. The observed sample mean is above 50 mm, but that difference is not especially unusual relative to its estimated standard error under the null model.
Two Sample Means: Distance in Either Direction
For a two-sample t test, the statistic measures how many estimated standard errors the observed difference in sample means is from the null difference. If the null says the population means are equal, the reference value is zero. A two-sided p-value counts results with a difference at least as far from zero as the observed difference, in either direction.
Worked Example: A Two-Sided Test of Battery Life
In a fictional comparison, technicians independently randomly sample eight batteries of Type A and eight of Type B from their respective production runs. Battery life is measured in hours. Type A has a sample mean of 23 hours and standard deviation of 3 hours; Type B has a sample mean of 20 hours and standard deviation of 3 hours. Each sample is less than 10% of its production run, and plots show no strong skewness or outliers. Test for a difference in population mean battery life at \(\alpha=0.05\).
Let \(\mu_1\) be the true mean battery life for Type A batteries from its production run, and \(\mu_2\) the true mean battery life for Type B batteries from its production run. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2\ne0\).
Use an unpooled two-sample t test. The batteries were randomly sampled, and the two samples contain different batteries, so there is no pairing and the groups are independent. The 10% condition holds for both samples. With only eight batteries per group, the plots matter; the stated lack of strong skewness or outliers supports the Nearly Normal condition.
The observed difference is \(23-20=3\) hours. Calculate its estimated standard error from the two sample variances separately, then standardize the difference from the null value of zero.
The Welch degrees of freedom are 14. For the two-sided alternative, the p-value includes both \(T_{14}\leq-2.00\) and \(T_{14}\geq2.00\). A t calculation gives \(p\approx0.06529\), rounded.
Because \(0.06529\) is greater than \(0.05\), fail to reject \(H_0\). These data do not provide convincing evidence of a difference in the true mean battery life of Type A and Type B batteries from their respective production runs.
The value \(0.06529\) is the probability, assuming equal population means, of a t statistic at least two units from zero in either direction. It is not the probability that the population means are equal, nor the probability that the observed difference of 3 hours occurred by chance in some general sense.
When the Observed Difference Points Toward the Alternative
The same sized t statistic can lead to different p-values depending on the alternative. This is why the research question determines the hypotheses before the data are examined. A one-sided test counts results in its specified direction; a two-sided test counts results in both directions.
Worked Example: A Directional Comparison of Study Times
A fictional school program compares weekly study time for students using a new planning tool and students using the usual method. The groups are independent random samples of eight students each. The new-tool group has a sample mean of 17 hours and standard deviation of 3 hours; the usual-method group has a sample mean of 20 hours and standard deviation of 3 hours. The samples are each less than 10% of their target populations, and group plots show no strong skewness or outliers. Researchers ask whether students using the new tool study fewer hours on average.
Let \(\mu_1\) and \(\mu_2\) be the true mean weekly study times for the new-tool and usual-method populations, respectively. The hypotheses are \(H_0:\mu_1-\mu_2=0\) and \(H_a:\mu_1-\mu_2<0\). An unpooled two-sample t test is appropriate: the groups are independent, the random samples support inference to their respective populations, the 10% condition holds for each, and the plots support the Nearly Normal condition for these small samples.
The observed difference in sample means is \(17-20=-3\) hours. The standard error is the same calculation used for two equal sample standard deviations and sample sizes:
Welch degrees of freedom are 14. Because the alternative points left, the p-value is the area at or below \(-2.00\), not the area in both tails. By symmetry, this is half the two-sided p-value from the preceding example: \(p\approx0.065288/2=0.032644\), or \(0.03264\) rounded. If the question instead asked whether the means differed in either direction, both tails would be included and the p-value would be \(0.06529\).
The calculation illustrates the meaning of “as extreme” precisely: for this left-tailed test, a result is at least as extreme as the observed one when its t statistic is \(-2.00\) or smaller. The alternative, not the fact that the statistic is negative by itself, defines which results count.
What a P-Value Does Not Say
- It is not the probability that the null hypothesis is true. The p-value is calculated assuming the null hypothesis is true; it does not assign a probability to the hypothesis itself.
- It is not the probability of getting the exact sample mean again. A p-value concerns a range of results at least as extreme as the observed statistic, not one exact value.
- It is not a measure of practical importance. A small p-value can occur for a small difference when the standard error is small. Consider the size and units of the difference as well as the test result.
- It does not repair a poor study design. The test conditions and the way the data were collected still matter. A small p-value does not remove bias or establish cause and effect in an observational study.
- It is not automatically two-sided. The alternative hypothesis determines whether one tail or both tails are counted.
A confidence interval and a p-value are related but not interchangeable. A confidence interval gives a range of plausible values for a population mean or difference in means. A p-value evaluates a particular null value against an alternative. As discussed in “Linking Two-Sample Intervals and Tests,” a two-sided test at significance level \(\alpha\) corresponds to checking whether the matching \(100(1-\alpha)\%\) interval includes the null value. The interval also gives information about the plausible size and direction of the difference.
Common Mistakes and AP Exam Tips
- Describing the p-value as the probability the null is true: Say, “Assuming the null hypothesis is true, the probability of a test statistic at least as extreme as the observed statistic is…”
- Forgetting the standard-error scale: A t statistic is not simply the difference in sample means. Show how the difference is divided by its estimated standard error.
- Using the wrong tail: Read \(H_a\) before calculating the p-value. For \(H_a:\mu_1-\mu_2<0\), use the left-tail area; for a two-sided alternative, include both tails.
- Calling a sample result “unlikely” without specifying the model: State that the probability is calculated assuming \(H_0\) is true and using the appropriate t distribution and degrees of freedom.
- Treating a large p-value as proof of equality: A large p-value means the data do not provide convincing evidence against \(H_0\) at the chosen significance level. It does not prove the population means are equal.
- Reporting only the calculator output: Include the observed sample result, standard error, t statistic, degrees of freedom, tail direction, and p-value. Then state what the result says about the population means in context.
A clear explanation connects the observed mean or mean difference to its t statistic, identifies which results count as at least as extreme, and describes the probability under the null model. Keep the distinction clear: the p-value describes how compatible the sample result is with the null hypothesis, not the probability that the null hypothesis is true.
Check Your Understanding
For each question, focus on the null model, the observed t statistic, and the alternative hypothesis.
- In your own words, define a p-value for a one-sample t test of a population mean.
- A two-sided t test produces \(t=1.8\). Which parts of the t distribution count as at least as extreme as the observed statistic?
- A right-tailed test has an observed statistic below zero. Should the p-value be the left-tail area or the right-tail area? Explain what determines the choice.
- Why is “the p-value is the probability that the null hypothesis is true” an incorrect interpretation?
- A two-sample test has \(H_a:\mu_1-\mu_2<0\) and \(t=-1.5\). Describe which tail area is the p-value and what assumption is used to calculate it.