Tutorials › AP Statistics › Parameters Versus Statistics for Proportions

Sampling distributions for proportions · Tutorial 401 of 1000

Parameters Versus Statistics for Proportions

Learn which proportion describes the whole population, which describes a sample, and how to calculate and interpret each one.

Intermediate 9 min read

What You'll Learn

  • Define the population proportion p for a specified group and characteristic.
  • Identify the sample proportion p-hat as a statistic calculated from sample data.
  • Calculate p-hat using the number of successes divided by the sample size.
  • Translate a reported sample percentage into a count when the information allows it.
  • Explain why a sample proportion does not automatically equal the population proportion.
  • Keep p and p-hat notation aligned with the population, sample, and time period in context.

Two Proportions That Describe Different Things

In Exam Practice on Evaluating Probability Models, you practiced defining a random variable and keeping a model’s claims separate from observed results. Proportions require a similar distinction. A proportion for an entire population is a parameter; a proportion calculated from a sample is a statistic. They may have similar numerical values, but they refer to different groups and play different roles.

Suppose a poll reports that 52% of 400 sampled voters favor a proposal. The 52% describes the sample: it is the fraction of those 400 people who said they favor the proposal. The proportion of all voters in the population who favor it is a different quantity. The sample result may provide information about that population quantity, but it is not automatically equal to it.

Definition: The population proportion \(p\) is the fraction of individuals in a specified population who have a particular categorical characteristic. The sample proportion \(\hat{p}\) is the fraction of individuals in a sample who have that characteristic. The symbol \(p\) names a parameter; \(\hat{p}\) names a statistic.

The characteristic being counted is often called a success, even when it is not a desirable outcome. For example, if the question is whether a voter favors a proposal, “favors the proposal” is the success category and “does not favor the proposal” is the other category. Define success before calculating so that the numerator is clear.

The population and the characteristic must both be specified. “The proportion who favor the proposal” is incomplete unless we know which people are meant, and possibly when their opinions are being described. For a particular population, characteristic, and time, \(p\) is one fixed value. It may be unknown, but unknown does not mean that the value changes from person to person. A different sample can produce a different \(\hat{p}\).

Notation and the Sample-Proportion Formula

Let \(x\) be the number of sampled individuals with the characteristic of interest, and let \(n\) be the total number of individuals in the sample. The sample proportion is the number in the success category divided by the sample size. It can be written as a decimal, fraction, or percentage, but calculations commonly use the decimal form.

$$ \hat{p}=\frac{x}{n} $$

The “hat” over \(p\) signals that the value is calculated from sample data. In contrast, \(p\) refers to the population proportion, not to the sample count or the sample’s percentage. If a question gives only a sample percentage and sample size, that percentage is \(\hat{p}\); the population proportion \(p\) is not thereby known.

SymbolWhat it describesHow it is obtained
\(p\)The proportion in the specified population with the characteristicA population parameter; often unknown
\(\hat{p}\)The proportion in the observed sample with the characteristic\(x/n\), using the sample’s success count and size
\(x\)The number of sampled individuals with the characteristicCount the successes in the sample
\(n\)The number of individuals in the sampleCount all sampled individuals in the group being analyzed

These symbols are connected, but they are not interchangeable. The population parameter \(p\) is the quantity a question about the whole population may ask about. The statistic \(\hat{p}\) is an observed summary of a particular sample. Later, when you study the sampling distribution of \(\hat{p}\), you will examine how this statistic behaves across repeated samples. For now, the essential distinction is what group each symbol describes.

Worked Example: Interpreting a Voter Poll

Worked Example: Interpreting a Voter Poll

A fictional poll asks 400 sampled voters in a district whether they favor a local transportation proposal. Exactly 52% of those sampled say yes. Identify \(x\), \(n\), and \(\hat{p}\), and distinguish the sample result from the population parameter.

Define the success. Let a success mean that a sampled voter says they favor the proposal. The sample size is \(n=400\). Because the reported percentage is exactly 52%, the number who favor the proposal is:

$$ x=0.52(400)=208 $$

Calculate the sample proportion. There are 208 successes among 400 sampled voters, so:

$$ \hat{p}=\frac{x}{n} =\frac{208}{400} =0.52 $$

Thus, \(\hat{p}=0.52\), or 52%. In context, 52% of the 400 voters in this sample said that they favor the proposal. Let \(p\) represent the proportion of all voters in the district who favor the proposal. That population proportion is the parameter of interest, and this poll’s information alone does not tell us its exact value.

The poll’s 52% is not \(p\) merely because it is the only percentage given. It is the observed sample statistic \(\hat{p}\). A separate sample of district voters could yield a different percentage. Also, this calculation identifies the statistic but does not by itself establish whether the sample represents all district voters well.

If a percentage is rounded rather than exact, multiplying it by \(n\) may not recover the exact count. For instance, a reported 52% could be a rounded summary. In that case, use the count if it is supplied; otherwise, describe the sample proportion as approximately 0.52 rather than claiming an exact success count.

Worked Example: A Sample Statistic and Its Population Parameter

Worked Example: A Sample Statistic and Its Population Parameter

A fictional environmental club takes a sample of 120 trees from a defined city park and records whether each tree has a particular leaf condition. Twenty-seven sampled trees have the condition. Let \(p\) be the proportion of all trees in that park with the condition, and calculate and interpret the sample proportion.

Identify the quantities. The characteristic of interest is having the leaf condition. In the sample, \(x=27\) trees have it, and the sample size is \(n=120\). The parameter \(p\) refers to all trees in the defined park; \(\hat{p}\) refers only to the 120 sampled trees.

Calculate.

$$ \hat{p} =\frac{x}{n} =\frac{27}{120} =0.225 $$

As a percentage, \(0.225(100)=22.5\%\). The sample proportion is 0.225, or 22.5%. In context, 22.5% of the sampled trees had the leaf condition. The proportion of all trees in the park with the condition is \(p\); it is not established as exactly 22.5% just by observing this sample.

Notice that both proportions refer to the same characteristic but to different groups. If the wording changes the target group—for example, from all trees in this park to all trees in the city—the parameter being described changes too. A clear definition of \(p\) prevents that shift from going unnoticed.

Comparing Statistics from Two Samples

If two samples are taken from the same population, their sample proportions can differ. Each one is a statistic calculated from its own sample. The population parameter remains the proportion for the specified population and characteristic, while the two observed statistics summarize separate groups of sampled individuals. This difference in notation is useful even before studying why sample proportions vary.

Worked Example: Two Samples of Community Gardeners

A fictional survey asks community gardeners whether they compost food scraps. One sample includes 80 gardeners, of whom 37 compost. A second sample includes 120 gardeners, of whom 54 compost. For each sample, calculate the sample proportion and identify the population parameter the surveys aim to describe.

Define the characteristic. A success is a gardener who composts food scraps. Let \(p\) be the proportion of all gardeners in the defined community who compost. For the first sample, \(x_1=37\) and \(n_1=80\):

$$ \hat{p}_1=\frac{37}{80}=0.4625 $$

The first sample proportion is 0.4625, or 46.25%. For the second sample, \(x_2=54\) and \(n_2=120\):

$$ \hat{p}_2=\frac{54}{120}=0.45 $$

The second sample proportion is 0.45, or 45%. Each value describes its own sample. Both surveys aim to provide information about the same parameter \(p\), assuming the target population and the definition of “composts food scraps” are the same.

The two statistics are not identical: 0.4625 differs from 0.45. That does not mean there are two population parameters for the same population and characteristic. Nor do these calculations alone establish which statistic is closer to \(p\). They simply report what was observed in each sample. Questions about how much sample statistics tend to vary, and what that variation can tell us about \(p\), come next.

How to Keep the Symbols Straight

When reading a question, first identify the population and the characteristic. Then identify whether the information describes the entire population or only a sample. This sequence makes it easier to assign \(p\) and \(\hat{p}\) correctly.

1
Name the population.
Be precise about which individuals the parameter describes, such as all registered voters in a district or all trees in one park.
2
Define the characteristic.
State which category counts as a success, such as favoring a proposal or having a particular leaf condition.
3
Classify the quantity.
Use \(p\) for the proportion in the population and \(\hat{p}\) for the proportion in a sample.
4
Calculate only from sample data.
For a sample, divide the success count \(x\) by the sample size \(n\): \(\hat{p}=x/n\).

A sample statistic is a number calculated from observed data. It can be known exactly once the sample has been observed, even when the population parameter remains unknown. Conversely, a population parameter may be fixed and meaningful even if no one has measured the entire population. These are separate ideas: whether a quantity is known and whether it describes a population or a sample.

Common Mistakes and AP Exam Communication

A full-credit response makes clear which group a proportion describes. A number without that group can be ambiguous, especially when a question mentions both a survey sample and a larger population.

  • Calling a sample percentage \(p\). If the percentage comes from the observed sample, label it \(\hat{p}\). For example, 52% of 400 sampled voters is \(\hat{p}=0.52\), not automatically the district’s \(p\).
  • Using the sample size as the success count. In \(\hat{p}=x/n\), \(n\) counts everyone in the sample and \(x\) counts only those with the defined characteristic. Confirm that the numerator is a count and the denominator is the full sample size.
  • Leaving success undefined. “Success” is a label for the category being counted, not a judgment that the outcome is good. Say exactly what qualifies.
  • Assuming \(p=\hat{p}\). A sample proportion may be used to learn about a population proportion, but the two refer to different groups. Do not state that they are equal unless the problem explicitly gives the population value or the entire population has been measured.
  • Changing the population without changing the parameter definition. A proportion for one district is not automatically the proportion for a whole state. Name the target population each time.
  • Confusing a decimal with a percent. The sample proportion 0.52 is the same amount as 52%. Keep the scale consistent in calculations and use words to clarify the context.
AP Exam Tip: Define \(p\) in words for the entire target population, define the sample’s success category, and show \(\hat{p}=x/n\) when calculating from sample data. Then interpret the statistic as a proportion of the sampled individuals, not as a guaranteed exact description of the population.

Key Takeaway

The distinction is about what the number describes, not merely how the number is written. \(p\) is the population parameter for a specified characteristic; \(\hat{p}\) is the sample statistic calculated by dividing the sample’s success count by its size. A sample statistic may inform us about a population parameter, but it is not the parameter itself.

Key takeaway: Use \(p\) for the proportion in the population and \(\hat{p}=x/n\) for the proportion observed in a sample. Always name the population, sample, and characteristic so the notation has a clear meaning.

Check Your Understanding

For each question, identify the population or sample described and use \(p\) and \(\hat{p}\) appropriately.

  1. A survey of 250 library visitors finds that 175 prefer digital notices. Define a success, calculate \(\hat{p}\), and state what population proportion \(p\) might represent if the target is all visitors to that library during a specified month.
  2. A report says that 38% of a sample of customers chose delivery. Is 0.38 \(p\) or \(\hat{p}\)? Explain what group it describes.
  3. In a sample of 90 devices, 9 have a particular defect. Identify \(x\) and \(n\), calculate \(\hat{p}\), and express the result as a percentage.
  4. Two samples from the same defined population produce \(\hat{p}_1=0.41\) and \(\hat{p}_2=0.46\) for the same characteristic. Do these values mean there are two population parameters? Explain.
  5. Why can the value of \(p\) be unknown even though it is a fixed parameter for a specified population and characteristic?