Tutorials › AP Statistics › Mean of the Sampling Distribution of p-hat

Sampling distributions for proportions · Tutorial 404 of 1000

Mean of the Sampling Distribution of p-hat

See why the long-run mean of sample proportions equals the population proportion, and practice calculating and interpreting that mean.

Intermediate 9 min read

What You'll Learn

  • Identify the mean of the sampling distribution of \(\hat{p}\)
  • Explain why \(\hat{p}\) is an unbiased estimator of \(p\)
  • Connect the expected number of successes to the mean of \(\hat{p}\)
  • Calculate the mean when \(p=0.40\) and when \(p=0.65\)
  • Distinguish the sampling-distribution mean from the result of one sample or one simulation

Why the Sampling Distribution Has a Center

In What a Sampling Distribution of p-hat Is, you learned that \(\hat{p}\) varies from sample to sample. In Simulating Sample Proportions with Repeated Samples, you saw that a plot of many simulated values can reveal where those values cluster. Now we will identify the exact mean of the sampling distribution: under the sampling process described here, it equals the population proportion \(p\).

This is a statement about the distribution over all possible samples of the same size, with each sample weighted according to its chance of being selected. It does not say that every sample proportion equals \(p\). A particular sample may have a proportion above or below \(p\); the claim concerns the long-run center across repeated samples.

Definition: The mean of the sampling distribution of \(\hat{p}\), written \(\mu_{\hat{p}}\), is the expected value of the sample proportion over the possible samples produced by the stated sampling process. For a random sample from a population with proportion \(p\), \(\mu_{\hat{p}}=p\).

Why the Mean Equals \(p\)

Suppose a population has \(N\) individuals, of whom \(K\) have a characteristic called a success. The population proportion is \(p=K/N\). A random sample of fixed size \(n\) produces a success count \(X\), and the sample proportion is \(\hat{p}=X/n\).

In a simple random sample, each population member has the same chance, \(n/N\), of being selected. So, among the \(K\) population members who are successes, the expected number selected is \(K(n/N)\). Because \(K/N=p\), this expected count is \(np\). The sample proportion is the success count divided by \(n\), so its mean is the expected success count divided by \(n\):

$$ \mu_{\hat{p}}=\frac{np}{n}=p $$

The same reasoning applies to repeated trials in which each trial has success probability \(p\): the expected number of successes in \(n\) trials is \(np\), and dividing that count by \(n\) gives a mean sample proportion of \(p\). The important idea is that the sample proportion is based on a count of successes, and the average count per sample is \(np\).

Formula: For a random sample of fixed size \(n\) from a population with proportion \(p\), the mean of the sampling distribution of the sample proportion is:
$$ \mu_{\hat{p}}=p $$

This result concerns the mean; it does not describe how much the sample proportions vary. As established in Normal Models for Sample Means and Proportions, the standard error describes their spread. The next tutorial develops the formula for that spread. For now, keep the roles distinct: \(p\) gives the center, while the standard error concerns variability around that center.

Unbiasedness: A Property of the Sampling Method

A statistic is an unbiased estimator of a parameter if the mean of its sampling distribution equals that parameter. Since \(\mu_{\hat{p}}=p\), the sample proportion \(\hat{p}\) is an unbiased estimator of the population proportion \(p\) under the stated random-sampling process.

“Unbiased” does not mean that an estimate from one sample must be correct, or that it will equal the parameter exactly. It means that if we repeatedly take samples in the same way and average their sample proportions, that average is centered at the population proportion. Individual sample proportions can still differ from \(p\).

The simulation in the preceding tutorial illustrated this distinction. Its plotted sample proportions clustered near the model value \(p=0.40\), but the center of one finite simulation need not equal 0.40 exactly. A simulation produces a limited set of outcomes; the mean of the theoretical sampling distribution describes all possible outcomes under the model.

Definition: An estimator is unbiased for a parameter when the mean of its sampling distribution equals that parameter. Thus, \(\hat{p}\) is unbiased for \(p\) because \(\mu_{\hat{p}}=p\). Unbiasedness describes the long-run center, not the accuracy of every individual estimate.

Worked Example: Find the Mean When \(p=0.40\)

Worked Example: Find the Mean When \(p=0.40\)

A school district’s student population has a proportion \(p=0.40\) who bring a reusable water bottle to school. A simple random sample of \(n=25\) students is selected. Find the mean of the sampling distribution of \(\hat{p}\), and explain the result.

Find the expected number of successes. Here, a success is a sampled student who brings a reusable bottle. The expected count is:

$$ np=25(0.40)=10 $$

Convert the expected count to a sample proportion. The sample proportion divides the success count by 25, so its mean is:

$$ \mu_{\hat{p}}=\frac{10}{25}=0.40 $$

Therefore, the sampling distribution of the proportion of sampled students who bring a reusable bottle has mean 0.40. Across all possible random samples of 25 students, the sample proportions are centered at 0.40. A single sample might yield, for example, \(\hat{p}=0.36\) or \(\hat{p}=0.48\); that variation does not change the mean.

Worked Example: Find the Mean When \(p=0.65\)

Worked Example: Find the Mean When \(p=0.65\)

A population of garden seedlings has a proportion \(p=0.65\) that survive a specified growing period. A random sample of \(n=40\) seedlings is observed. Find the mean of the sampling distribution of \(\hat{p}\).

Calculate the expected number of surviving seedlings in a sample.

$$ np=40(0.65)=26 $$

Divide by the sample size to get the mean sample proportion.

$$ \mu_{\hat{p}}=\frac{26}{40}=0.65 $$

The mean of the sampling distribution is 0.65. In repeated random samples of 40 seedlings, the sample proportion that survive is centered at the population proportion of 0.65. The expected count of 26 is an average across samples, not a promise that every sample will contain exactly 26 surviving seedlings.

Worked Example: Verify the Mean from Possible Outcomes

Worked Example: Verify the Mean from Possible Outcomes

Imagine two independent trials, each with success probability \(p=0.65\). Let \(X\) be the number of successes, so \(\hat{p}=X/2\). Verify the mean by calculating the probabilities of the possible sample proportions and using them to find the expected value.

The possible success counts are 0, 1, and 2. Their probabilities are:

$$ P(X=0)=(0.35)^2=0.1225,\quad P(X=1)=2(0.65)(0.35)=0.4550,\quad P(X=2)=(0.65)^2=0.4225 $$

These probabilities add to 1: \(0.1225+0.4550+0.4225=1.0000\). The corresponding sample proportions are 0, 0.50, and 1.00. Multiply each possible value by its probability and add:

$$ \mu_{\hat{p}} =(0)(0.1225)+(0.50)(0.4550)+(1.00)(0.4225) =0+0.2275+0.4225 =0.6500 $$

The mean is 0.65, matching \(p\). Notice that 0.65 is not one of the possible sample proportions for \(n=2\), yet it is their probability-weighted mean. The mean of a sampling distribution need not be a value that an individual sample can produce.

Conditions and What the Result Does—and Does Not—Say

The formula applies when the sampling process gives each sampled observation the population proportion \(p\) as its success probability, or when the sample is a simple random sample from the population whose proportion is \(p\). The sample size is fixed. Random selection matters: a biased selection process might systematically overrepresent or underrepresent people with the characteristic, so its sample proportions need not be centered at the population proportion.

For a simple random sample without replacement, the mean result is exact. The 10% condition is often used when treating observations from a finite population as approximately independent for other calculations. It is not needed to establish this mean: each individual’s chance of selection is \(n/N\), which gives the expected success count \(np\) even when sampling without replacement.

The equality \(\mu_{\hat{p}}=p\) does not say that \(\hat{p}\) is normally distributed, nor does it determine the probability of any particular sample proportion. It identifies only the center. To answer questions about spread or probabilities, we need additional information and conditions, developed in later tutorials.

Conditions: Interpret \(\mu_{\hat{p}}=p\) for a fixed sample size and a random sampling process that represents the population or model of interest. Biased selection can shift the center. The mean result itself does not require the 10% condition for a simple random sample.

Common Mistakes and AP Exam Communication

A complete response names the statistic and parameter, states the mean relationship, and interprets it in context. If the question gives \(n\), showing the expected count \(np\) and dividing by \(n\) makes the connection clear.

  • Claiming every sample proportion equals \(p\). Unbiasedness is about the long-run mean, not every sample. Say the sample proportions are centered at \(p\), not that each sample produces \(\hat{p}=p\).
  • Confusing \(p\) and \(\hat{p}\). The population proportion \(p\) is a parameter; \(\hat{p}\) is a statistic that changes from sample to sample. Identify which one describes the population and which one describes a sample.
  • Calling \(np\) the mean of \(\hat{p}\). The value \(np\) is the expected success count. Divide by \(n\) to obtain the mean sample proportion: \(np/n=p\).
  • Treating an observed simulation average as the exact mean. A finite run can have an average above or below \(p\). Use “in these simulated samples” for its observed average, and distinguish that from the theoretical mean.
  • Assuming the mean gives the spread or shape. The fact that \(\mu_{\hat{p}}=p\) tells where the sampling distribution is centered. It does not establish a standard deviation, normal shape, or tail probability.
AP Exam Tip: Write, “The sampling distribution of \(\hat{p}\) is centered at the population proportion \(p\), so \(\hat{p}\) is an unbiased estimator of \(p\).” Then name the characteristic and population in context. If you calculate \(np\), identify it as the expected number of successes, not the mean of \(\hat{p}\).

Key Takeaway

The mean of the sampling distribution of \(\hat{p}\) equals the population proportion \(p\). For \(p=0.40\), the mean is 0.40; for \(p=0.65\), the mean is 0.65. This makes \(\hat{p}\) an unbiased estimator of \(p\): repeated random samples are centered at the population value, even though individual sample proportions vary.

Key takeaway: The expected success count is \(np\), and dividing by the sample size gives \(\mu_{\hat{p}}=np/n=p\). Unbiasedness describes the sampling distribution’s center, not a guarantee about any one sample.

Check Your Understanding

Answer each question using the relationship between the population proportion and the mean of the sampling distribution.

  1. A population has \(p=0.40\), and random samples have size \(n=50\). What is the expected number of successes, and what is \(\mu_{\hat{p}}\)?
  2. A sampling distribution of \(\hat{p}\) has mean 0.65. What is the population proportion \(p\), and why is \(\hat{p}\) unbiased for it?
  3. In Worked Example 3, why can the mean be 0.65 even though the only possible sample proportions are 0, 0.50, and 1.00?
  4. A student says, “Because \(\hat{p}\) is unbiased, every sample proportion will equal \(p\).” Explain the error.
  5. Does a simulation average of 0.42 from samples generated under \(p=0.40\) contradict \(\mu_{\hat{p}}=p\)? Explain.