Tutorials › AP Statistics › Finding a P-Value from a z Statistic with normalcdf

One-proportion hypothesis tests · Tutorial 468 of 1000

Finding a P-Value from a z Statistic with normalcdf

Use the direction of the alternative hypothesis to choose the correct normalcdf tail, then double one tail only when the test is two-sided.

Intermediate 10 min read

What You'll Learn

  • Explain what a p-value measures under the null hypothesis
  • Match a less-than or greater-than alternative to the correct normalcdf tail
  • Calculate a two-sided p-value by doubling the smaller tail
  • Enter normalcdf bounds and standard Normal parameters correctly
  • Interpret p-values and avoid confusing them with probabilities that the null hypothesis is true

From a z Statistic to a Tail Probability

In Why Tests Use \(p_0\) in the Standard Deviation, you saw why a one-proportion \(z\)-test measures the sample result against a model that assumes the null hypothesis is true. Once the test statistic \(z\) has been calculated, the next step is to find the probability, under that null model, of getting a result at least as extreme as the one observed.

That probability is the p-value. For a one-proportion \(z\)-test that meets its conditions, the test statistic is compared with a standard Normal distribution centered at 0. The alternative hypothesis determines which part, or parts, of the curve count as “at least as extreme.” Therefore, you must identify the direction of \(H_a\) before using normalcdf.

Definition: A p-value is the probability, assuming the null hypothesis is true, of obtaining a test statistic at least as extreme as the observed statistic in the direction specified by the alternative hypothesis.

Use the Alternative Hypothesis to Choose the Tail

The sign of \(z\) tells you whether the observed sample proportion is above or below \(p_0\). The alternative hypothesis tells you which direction counts as evidence against \(H_0\). Match those directions when you select a tail.

  • For \(H_a:p>p_0\), evidence is in the upper direction. Find the area to the right of the observed \(z\).
  • For \(H_a:p<p_0\), evidence is in the lower direction. Find the area to the left of the observed \(z\).
  • For \(H_a:p\ne p_0\), departures in either direction count. Find the area beyond the observed distance from 0 in both tails.

On a TI-84, normalcdf(lower bound, upper bound, mean, standard deviation) calculates the area under a Normal curve between the bounds. For a standard Normal distribution, use a mean of 0 and a standard deviation of 1. Because the curve extends indefinitely, calculator work commonly uses \(-1\text{E}99\) and \(1\text{E}99\) as practical substitutes for negative and positive infinity.

$$ \begin{aligned} H_a:p>p_0 &: \quad P(Z\geq z_{\text{obs}})=\operatorname{normalcdf}(z_{\text{obs}},1\text{E}99,0,1)\\ H_a:p<p_0 &: \quad P(Z\leq z_{\text{obs}})=\operatorname{normalcdf}(-1\text{E}99,z_{\text{obs}},0,1) \end{aligned} $$

Here, \(Z\) is the standard Normal test statistic under the null model, and \(z_{\text{obs}}\) is the test statistic calculated from the sample. The calculator’s use of “less than or equal to” or “greater than or equal to” does not change the area in this setting: a Normal distribution is continuous, so the probability of exactly one value is 0.

Two-Sided Tests: Double the Tail Area

For \(H_a:p\ne p_0\), a result far below \(p_0\) can be evidence against the null, just as a result equally far above \(p_0\) can be. The two-sided p-value includes both tails. Because the standard Normal curve is symmetric around 0, you can find the area beyond \(\lvert z_{\text{obs}}\rvert\) in one tail and double it.

$$ \text{Two-sided p-value} =2P(Z\geq |z_{\text{obs}}|) =2\operatorname{normalcdf}(|z_{\text{obs}}|,1\text{E}99,0,1) $$

Equivalently, if the observed statistic is negative, find the area to its left and double it. If it is positive, find the area to its right and double it. The absolute value gives the distance from 0 without regard to direction. Do not double a one-sided p-value: doing so would count results in a direction that the alternative hypothesis does not identify as evidence.

Key distinction: A one-sided p-value is the area in the tail specified by \(H_a\). A two-sided p-value is the combined area in both tails beyond the observed distance from 0. “Double the tail” applies to a two-sided test, not automatically to every p-value.

Worked Examples

Worked Example: An Upper-Tail Test

A fictional random sample of 100 library cardholders includes 60 who used the library’s digital-book app last month. Test \(H_0:p=0.50\) against \(H_a:p>0.50\), where \(p\) is the true proportion of cardholders who used the app last month. Find the p-value and state the test conclusion at \(\alpha=0.05\).

State: The parameter is the true proportion of library cardholders who used the digital-book app last month. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\).

Plan and check conditions: The cardholders were randomly sampled, so the Random condition is met. The sample was drawn without replacement from 2,400 cardholders, and \(100\leq0.10(2400)=240\), so the 10% condition is met. Under \(H_0\), the expected numbers of successes and failures are \(np_0=100(0.50)=50\) and \(n(1-p_0)=100(0.50)=50\). Both are at least 10, so the test’s Large Counts condition is met. A one-proportion \(z\)-test is appropriate.

Do: The sample proportion is \(60/100=0.60\). The null standard error is \(\sqrt{0.50(0.50)/100}=0.05\), so the test statistic is:

$$ z=\frac{0.60-0.50}{0.05}=2.00 $$

Because the alternative is greater than, use the area to the right of 2.00:

$$ \text{p-value} =\operatorname{normalcdf}(2.00,1\text{E}99,0,1) \approx 0.0228 $$

Conclude: If the true proportion of cardholders who used the app were 0.50, the probability of obtaining a test statistic of 2.00 or greater would be about 0.0228. Since \(0.0228<0.05\), reject \(H_0\). The sample provides convincing evidence that more than 50% of library cardholders used the digital-book app last month.

Worked Example: A Lower-Tail Test

A fictional random sample of 100 community gardeners includes 40 who use drip irrigation. A planner tests \(H_0:p=0.50\) against \(H_a:p<0.50\), where \(p\) is the true proportion of community gardeners who use drip irrigation. The test statistic is \(z=-2.00\). Find and interpret the p-value.

The gardeners were randomly sampled from a group of 1,800, so the Random condition is met and \(100\leq0.10(1800)=180\) verifies the 10% condition. Under the null, the expected counts are \(100(0.50)=50\) successes and \(100(0.50)=50\) failures; both are at least 10. The conditions for the one-proportion \(z\)-test are met.

Because \(H_a\) says the proportion is lower, the relevant results lie to the left of the observed statistic. Enter the observed negative \(z\), not its absolute value:

$$ \text{p-value} =\operatorname{normalcdf}(-1\text{E}99,-2.00,0,1) \approx 0.0228 $$

Assuming the true proportion of gardeners who use drip irrigation is 0.50, there is about a 0.0228 probability of obtaining a test statistic of \(-2.00\) or less. The result is in the direction specified by the alternative. Notice that the area is the same as in the upper-tail example: the standard Normal curve is symmetric, and both observed statistics are 2 standard deviations from 0. The direction still matters, because each example uses a different alternative hypothesis.

Worked Example: A Two-Sided Test

A researcher tests whether a population proportion differs from a proposed benchmark. The hypotheses are \(H_0:p=p_0\) and \(H_a:p\ne p_0\), and the one-proportion test statistic is \(z=2.05\). Find the two-sided p-value. Treat the statistic as the value already calculated from the sample and the null standard error.

A two-sided alternative counts results at least as far from 0 in either direction. First find the upper-tail area beyond 2.05, then double it:

$$ \begin{aligned} \text{one-tail area} &=\operatorname{normalcdf}(2.05,1\text{E}99,0,1)\\ &\approx 0.0202\\ \text{two-sided p-value} &=2(0.0202)=0.0404 \end{aligned} $$

The same calculation can be viewed as adding the two tail areas. By symmetry, the lower-tail area below \(-2.05\) is also about 0.0202, so the combined probability is \(0.0202+0.0202=0.0404\). Thus, assuming \(H_0\) is true, the probability of a test statistic at least 2.05 units from 0 in either direction is about 0.0404.

Do not report 0.0202 as the p-value for this test: that is only one of the two relevant tails. Also do not use the area between \(-2.05\) and 2.05; that is the probability of statistics less extreme than the observed one.

Common Mistakes and What Full Credit Says

Finding the area is usually straightforward once the tail is clear. Many errors happen before the calculator is used: students may follow the sign of \(z\) instead of the alternative, or double an area when the test is one-sided. Check the hypotheses first, then connect the calculation to the context.

  • Choosing the tail from the sign of \(z\) alone. The sign shows where the observed result is; \(H_a\) determines which direction is evidence against \(H_0\). For \(H_a:p>p_0\), use the upper tail even if a sample happened to produce a negative \(z\).
  • Using the wrong bounds. An upper-tail calculation uses the observed \(z\) as the lower bound and a very large upper bound. A lower-tail calculation uses a very negative lower bound and the observed \(z\) as the upper bound.
  • Doubling a one-sided p-value. A one-sided alternative specifies one direction, so use only that tail. Double a tail area for a two-sided alternative because departures in both directions count.
  • Forgetting the second tail in a two-sided test. The area beyond \(z_{\text{obs}}\) on just one side is not the full two-sided p-value. Double the tail beyond \(\lvert z_{\text{obs}}\rvert\), or add both tail areas.
  • Calling the p-value the probability that \(H_0\) is true. A p-value is calculated under the assumption that \(H_0\) is true. It measures how unusual the observed statistic, or a more extreme one as defined by \(H_a\), would be under that assumption.
AP Exam Tip: Write the alternative hypothesis, identify the tail in words, show the normalcdf bounds, and report the probability. For a two-sided test, explicitly show the doubling. Interpret the result conditionally: “Assuming \(H_0\) is true, the probability of obtaining a test statistic at least this extreme in the direction(s) specified by \(H_a\) is …”

Key Takeaway

A z statistic gives a location on the standard Normal curve; the alternative hypothesis tells you which area to measure. Use the upper tail for a greater-than alternative, the lower tail for a less-than alternative, and both tails for a not-equal-to alternative. Then interpret the resulting probability under the null model, in context.

Key takeaway: Match the normalcdf tail to \(H_a\). For a two-sided test, double the area beyond \(\lvert z_{\text{obs}}\rvert\); for a one-sided test, do not double. A p-value is a probability assuming \(H_0\) is true, not the probability that \(H_0\) is true.

Check Your Understanding

For each question, identify the appropriate tail or tails and calculate or describe the p-value as requested.

  1. For \(H_a:p>p_0\) and \(z=1.50\), write the normalcdf expression for the p-value.
  2. For \(H_a:p<p_0\) and \(z=-1.75\), write the normalcdf expression for the p-value.
  3. A two-sided test has \(z=-1.80\). Write a normalcdf expression that uses symmetry and explain why it is doubled.
  4. For a greater-than alternative, the observed \(z\) is \(-1.20\). Should the p-value use the lower tail because \(z\) is negative, or the upper tail because of the alternative? Explain.
  5. In words, interpret a p-value of 0.03 for a test of \(H_0:p=p_0\) against \(H_a:p\ne p_0\). What does the 0.03 not mean?