Tutorials › AP Statistics › What a Sampling Distribution of p-hat Is

Sampling distributions for proportions · Tutorial 402 of 1000

What a Sampling Distribution of p-hat Is

See how the sample proportion varies across repeated samples of the same size, and how a simulated distribution makes that variation visible.

Intermediate 9 min read

What You'll Learn

  • Define the sampling distribution of the sample proportion.
  • Distinguish a sampling distribution from the distribution of individuals in one sample.
  • Track how success counts determine possible sample proportions when the sample size is fixed.
  • Read and interpret simulated results from repeated samples of size 50 when the population proportion is 0.30.
  • Describe the center and variability of the sampling distribution in context.

One Sample Proportion, Many Possible Results

In Parameters Versus Statistics for Proportions, you learned that \(p\) describes a population and \(\hat{p}=x/n\) describes a sample. Now imagine taking many random samples of the same size from a population with a known proportion \(p=0.30\). Each sample will have its own success count and its own \(\hat{p}\). The resulting values will generally vary, even though the population and sample size stay the same.

For example, suppose success means that a randomly selected item has a particular feature. In one sample of 50 items, 14 might have the feature, giving \(\hat{p}=14/50=0.28\). Another sample of 50 might have 17 successes, giving \(\hat{p}=17/50=0.34\). The population proportion has not changed; the sample results have.

Definition: The sampling distribution of \(\hat{p}\) is the distribution of the sample-proportion values from all possible random samples of the same size, drawn using the same sampling process from a population with a specified proportion \(p\). It describes how \(\hat{p}\) varies from sample to sample.

The sampling distribution is not the distribution of the individuals in one sample. In one sample, we record each individual’s category, such as success or not success. The sampling distribution instead records the resulting \(\hat{p}\) from each repeated sample. Its observations are sample proportions.

In reality, we usually do not take every possible sample. A simulation can imitate the sampling process many times and show the variation that repeated samples can produce. The simulation does not make the sample results identical or guarantee that any one result will occur.

What Varies and What Stays the Same?

To describe repeated samples, keep the population, characteristic, and sample size fixed. Let success mean that an individual has the characteristic of interest, and let \(p=0.30\) be the proportion of the population with that characteristic. For every sample, take \(n=50\) individuals and calculate the sample proportion.

Before a particular sample is selected, its success count is not yet known. We can represent the count from one sample by the random variable \(X\). The sample proportion from that sample is \(\hat{p}=X/50\). Across repetitions, \(X\) and \(\hat{p}\) can take different values. The population proportion \(p=0.30\), however, stays fixed in this setup.

$$ \hat{p}=\frac{X}{50} $$

Because the sample size is 50, the possible values of \(\hat{p}\) occur in steps of \(1/50=0.02\): 0, 0.02, 0.04, and so on up to 1. For instance, 14 successes produce \(14/50=0.28\), while 15 successes produce \(15/50=0.30\). The sampling distribution is therefore a distribution of discrete possible proportions, not a smooth list of every decimal between 0 and 1.

The sample size matters to the set of possible values. With 50 observations, each additional success changes \(\hat{p}\) by \(1/50=0.02\). But the central idea is not the size of that step: it is that the statistic is recalculated for each repeated sample, creating a distribution of results.

Worked Example: Simulating 1,000 Samples

Worked Example: Simulating 1,000 Samples

Imagine a population in which 30% of individuals have a particular characteristic. A computer simulation imitates 1,000 random samples, each of size 50. For each simulated individual, the computer assigns success with probability 0.30 and not success with probability 0.70. Within each repetition, it records the number of successes and divides by 50 to obtain \(\hat{p}\).

One way to implement the simulation is to generate a random integer from 00 to 99 for each simulated individual. Values 00 through 29 represent success, giving 30 of the 100 equally likely values. Values 30 through 99 represent not success, giving 70 of the 100 values. Generate 50 such outcomes for one sample, calculate its \(\hat{p}\), and repeat the whole process 1,000 times.

The following table shows one illustrative set of results. It groups the simulated samples by their success counts. The proportions in the second column are the corresponding ranges of \(\hat{p}\), since every sample has size 50.

Success count in a sampleCorresponding \(\hat{p}\) valuesNumber of simulated samplesPercent of 1,000 samples
0–80–0.16202%
9–110.18–0.2214014%
12–140.24–0.2835035%
15–170.30–0.3433033%
18–200.36–0.4013513.5%
21–500.42–1.00252.5%
Total—1,000100%

For example, the 350 samples with 12–14 successes have sample proportions from \(12/50=0.24\) through \(14/50=0.28\). Their share of all simulated samples is calculated as:

$$ \frac{350}{1000}=0.35=35\% $$

The number of simulated samples in the table adds to 1,000: \(20+140+350+330+135+25=1{,}000\). The listed percentages also add to 100%: \(2+14+35+33+13.5+2.5=100\%\). This confirms that the table accounts for all simulated repetitions.

The results gather most often near 0.30. They do not all equal 0.30: some are below it and some are above it. That is the variation the sampling distribution represents. This single simulation is an illustration of the process, not a claim about an actual study or a promise that a new run would produce exactly the same table.

Center and Variability Across Repeated Samples

As established in Normal Models for Sample Means and Proportions, the mean of the sampling distribution of \(\hat{p}\) is \(p\), and its standard deviation, called the standard error, is \(\sqrt{p(1-p)/n}\) when the sampling conditions apply. These describe the center and spread of the statistic over repeated samples. They do not describe the spread of individual responses within one sample.

For the setup with \(p=0.30\) and \(n=50\), the mean is 0.30. The standard error is:

$$ \sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.30(0.70)}{50}} =\sqrt{0.0042} \approx 0.0648 $$

A second way to check the standard-error calculation is to consider the success count: its standard deviation under this model is \(\sqrt{50(0.30)(0.70)}=\sqrt{10.5}\approx3.2404\) successes. Dividing by 50 converts that spread in counts to the spread in proportions: \(3.2404/50\approx0.0648\). The two calculations agree.

In context, the standard deviation of the sampling distribution is about 0.0648, or 6.48 percentage points. It describes the typical distance of a sample proportion from the sampling distribution’s mean of 0.30. It is not a guarantee that every sample proportion will be within 0.0648 of 0.30.

Key distinction: The population proportion \(p=0.30\) describes individuals in the population. The mean and standard deviation of the sampling distribution describe the behavior of \(\hat{p}\) across repeated samples. The standard error measures sample-to-sample variability.

Worked Example: Interpreting the Center and Spread

Worked Example: Interpreting the Center and Spread

Suppose a student says, “With \(p=0.30\) and samples of 50, the sample proportion will always be 0.30, because the sampling distribution is centered at 0.30.” Explain what is right and wrong in this statement.

Identify what the center describes. The mean of the sampling distribution is \(p=0.30\). This means that over all possible samples of size 50 under the stated process, the sample proportions are centered at 0.30. It does not say that every individual sample produces that value.

Check possible sample results. If a sample has 14 successes, its sample proportion is:

$$ \hat{p}=\frac{14}{50}=0.28 $$

If another sample has 17 successes, its sample proportion is:

$$ \hat{p}=\frac{17}{50}=0.34 $$

Both values differ from 0.30, and both are possible. The student is right that the sampling distribution is centered at 0.30, but wrong to conclude that each sample proportion must equal its center. In context, 0.30 is the mean sample proportion over repeated samples, not a required result for every sample.

Reading a Simulated Sampling Distribution

A simulation table or graph helps you see how often different sample-proportion values occurred in the simulated repetitions. For a graph, each dot or bar represents simulated samples with a particular \(\hat{p}\) or range of \(\hat{p}\) values. A dense region indicates values that appeared more often in that run; it does not change the population parameter.

The simulation results in the first example can also be used to describe a specific event. For instance, the table reports 25 simulated samples with at least 21 successes, which corresponds to \(\hat{p}\geq0.42\). The fraction of simulated results meeting that rule is:

$$ \frac{25}{1000}=0.025=2.5\% $$

This says that 2.5% of the simulated samples in that run had \(\hat{p}\geq0.42\). It does not say that the population proportion is 0.42, nor that exactly 2.5% of every possible set of 1,000 samples must meet the rule. As discussed in Interpreting Simulation Results Against a Model, a simulation proportion summarizes the results of the simulated repetitions under the stated model.

Worked Example: A Large Sample Proportion Is Still a Statistic

Worked Example: A Large Sample Proportion Is Still a Statistic

In one sample of 50 individuals from the population with \(p=0.30\), suppose 21 individuals have the characteristic. A student reports that the population proportion must be 0.42 because the sample proportion is 0.42. Use the simulation results to explain the distinction.

Calculate the observed sample proportion. Here, \(x=21\) and \(n=50\), so:

$$ \hat{p}=\frac{x}{n}=\frac{21}{50}=0.42 $$

Thus, 42% of the individuals in this particular sample had the characteristic. The value 0.42 is the sample statistic; the population parameter remains \(p=0.30\) in the setup. A sample statistic can differ from the population proportion because samples vary.

Relate the result to the simulation. In the illustrated run, 25 out of 1,000 simulated samples had at least 21 successes, or \(\hat{p}\geq0.42\). The arithmetic is \(25/1000=0.025=2.5\%\). So a sample proportion of 0.42 or larger appeared infrequently in that simulation, but it did appear. Its occurrence does not make \(p\) equal to 0.42; it is one possible sample result under the model.

The simulation does not prove that the model is correct or that the actual population has \(p=0.30\). It shows what repeated samples can look like when the simulation uses that stated population proportion and sampling process.

Common Mistakes and AP Exam Communication

When describing a sampling distribution, make clear what is being repeated and what the results represent. A full-credit explanation names the sample size, the statistic, and the population proportion or model that stays fixed.

  • Confusing the sample distribution with the sampling distribution. A sample distribution describes individual observations in one sample. A sampling distribution describes the statistic \(\hat{p}\) across repeated samples.
  • Treating \(p\) and \(\hat{p}\) as interchangeable. In this setup, \(p=0.30\) is fixed for the population, while \(\hat{p}\) varies from sample to sample. A sample proportion of 0.42 is not automatically the population proportion.
  • Claiming every sample proportion equals the center. The mean of the sampling distribution is \(p\), but individual sample proportions can fall above or below that mean.
  • Forgetting that the sample size is held fixed. A sampling distribution for samples of size 50 is not the same setup as one for a different sample size. State \(n\) when defining the distribution.
  • Interpreting standard error as a limit. A standard error describes typical variability; it does not give the largest possible distance from \(p\). Sample proportions can fall farther away.
  • Calling a simulated percentage an exact guarantee. A simulation estimates or illustrates behavior under its model. Different simulation runs can produce different counts and percentages.
AP Exam Tip: State that the sampling distribution is formed by calculating \(\hat{p}\) for repeated random samples of the same size from a population with proportion \(p\). Then describe its center or spread as properties of those sample proportions, not of the individuals in one sample.

Key Takeaway

A sample proportion is one result from one sample. The sampling distribution brings together the \(\hat{p}\) values from repeated samples with the same size and sampling process. For \(p=0.30\) and \(n=50\), the distribution is centered at 0.30, but its sample proportions vary around that center.

Key takeaway: Keep the population proportion \(p\) distinct from the varying statistic \(\hat{p}\). A sampling distribution describes the possible sample proportions and how they vary across repeated samples of a fixed size.

Check Your Understanding

Use the setup \(p=0.30\) and \(n=50\) where relevant. Explain what each quantity or result describes.

  1. In one sample, 16 of 50 individuals have the characteristic. Calculate \(\hat{p}\) and interpret it in context.
  2. In your own words, distinguish the distribution of individuals in one sample from the sampling distribution of \(\hat{p}\).
  3. Why can two random samples of size 50 from the same population produce different sample proportions even though \(p\) stays fixed?
  4. What does it mean to say that the sampling distribution of \(\hat{p}\) is centered at 0.30? Does that require every sample proportion to equal 0.30?
  5. A simulation produces 40 sample proportions at or above 0.42 out of 2,000 repetitions. Calculate the simulated percentage and state what it describes.