From a Sample Proportion to a Standardized Distance
In Finding Percentiles of a Sample Proportion with invNorm, you used a Normal model to locate a sample-proportion cutoff at a specified percentile. Now consider a value of \(\hat{p}\) that has already been observed. A z-score tells us how far that sample proportion is from the population proportion, measured in standard deviations of the sampling distribution.
The comparison is not simply between two proportions. A difference of \(0.04\), for example, represents a different amount of sampling variation depending on the sample size and the population proportion. Dividing the difference by the standard deviation of \(\hat{p}\) puts it on a common scale.
As established in Standard Deviation of p-hat Formula, the standard deviation of the sampling distribution is \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\), when the sampling process supports that model. Use this standard deviation to standardize the observed result.
How to Calculate and Interpret the z-Score
First find the observed sample proportion from the sample count: \(\hat{p}=x/n\), where \(x\) is the number of sampled individuals with the characteristic. Then calculate \(\sigma_{\hat{p}}\) using the population proportion and sample size. Finally, subtract \(p\) from \(\hat{p}\) and divide by \(\sigma_{\hat{p}}\).
The numerator, \(\hat{p}-p\), is the observed difference in proportions. The denominator measures the typical spread of sample proportions under the stated sampling model. The quotient has no units: it expresses the difference in standard-deviation units. For instance, \(z=1.4\) means the observed sample proportion is 1.4 standard deviations above \(p\), not that it is 1.4 percentage points above \(p\).
Identify what the sample proportion measures, its observed value, and the model’s population proportion \(p\).
Verify randomness, the 10% condition when sampling without replacement, and the Large Counts condition if using a Normal model.
Use \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\), based on the model’s \(p\), not the observed \(\hat{p}\).
Calculate \((\hat{p}-p)/\sigma_{\hat{p}}\), then state the direction and distance in standard deviations.
The condition checks matter for interpretation. The formula gives a standardized distance when its inputs are appropriate. If randomness or independence assumptions are not reasonable, the model’s standard deviation may not describe the sampling process well. If the Large Counts condition is not met, a Normal approximation may not be appropriate; calculating a z-score does not fix that problem.
Worked Example: A Sample Proportion Above the Model Value
Worked Example: Residents Who Compost Food Scraps
Suppose 40% of residents in a town compost food scraps. A random sample of 100 residents is selected without replacement from the town’s 5,000 residents. In the sample, 47 residents compost food scraps. Find and interpret the z-score for the observed sample proportion.
State. Let \(\hat{p}\) be the proportion of sampled residents who compost food scraps. The population proportion in the model is \(p=0.40\). The observed sample proportion is \(\hat{p}=47/100=0.47\).
Plan and check conditions. The sample is random. Because it is drawn without replacement, check the 10% condition: \(100\leq0.10(5{,}000)=500\), so it is met. The expected number who compost is \(np=100(0.40)=40\), and the expected number who do not is \(n(1-p)=100(0.60)=60\). Both are at least 10, so the Large Counts condition supports a Normal model.
Do. Calculate the standard deviation of the sampling distribution, then divide the observed difference by that standard deviation:
As a check, the observed difference is \(0.07\), and \(0.07/0.04899\approx1.429\), the same result.
Conclude. The observed sample proportion of residents who compost is about 1.43 standard deviations above the population proportion of 0.40, using the sampling-distribution model.
Reading the Sign and Size
The sign of the z-score comes from the order of subtraction: \(\hat{p}-p\). If the observed sample proportion exceeds the model proportion, the numerator is positive. If it is smaller, the numerator is negative. A result equal to \(p\) has z-score zero.
The absolute value \(|z|\) describes the distance from \(p\) in standard deviations, without giving the direction. For example, a z-score of \(-1.2\) is 1.2 standard deviations below \(p\), while \(+1.2\) is 1.2 standard deviations above \(p\). Keep both pieces when interpreting a signed score.
A z-score is not a probability. It reports a location on a standardized scale. When the Normal model is appropriate, standardizing also connects the value to the standard Normal model, whose mean is 0 and standard deviation is 1. But the z-score itself does not state the probability of the observed result or of results at least as far away.
Likewise, a z-score is not the raw difference in percentage points. If \(\hat{p}=0.47\) and \(p=0.40\), the raw difference is \(0.07\), or 7 percentage points. The z-score describes that difference relative to the standard deviation of sample proportions for the specified sample size and model.
Worked Example: A Sample Proportion Below the Model Value
Worked Example: Preference for a Local Bus Pass
In a large community, 18% of residents prefer a particular local bus pass. A random sample of 200 residents is selected without replacement from a population of 6,000. In the sample, 28 residents prefer the pass. Find and interpret the z-score for the sample proportion.
State. Let \(\hat{p}\) be the proportion of sampled residents who prefer the bus pass. The model uses \(p=0.18\), and the observed sample proportion is \(\hat{p}=28/200=0.14\).
Plan and check conditions. The sample is random. The 10% condition is met because \(200\leq0.10(6{,}000)=600\). The expected number who prefer the pass is \(np=200(0.18)=36\), and the expected number who do not is \(n(1-p)=200(0.82)=164\). Both expected counts are at least 10, so the Large Counts condition supports a Normal model.
Do. The standard deviation and z-score are:
Checking with the unrounded standard deviation gives \(-0.04/\sqrt{0.000738}\approx-1.472\), consistent with the reported value.
Conclude. The observed sample proportion of residents who prefer the bus pass is about 1.47 standard deviations below the population proportion of 0.18, according to the sampling-distribution model.
Why Sample Size Changes the Standardized Distance
The standard deviation \(\sigma_{\hat{p}}\) depends on both \(p\) and \(n\). For a fixed \(p\), a larger sample size makes \(\sigma_{\hat{p}}\) smaller, as explained in How Sample Size Changes the Spread of p-hat. Therefore, the same raw difference between \(\hat{p}\) and \(p\) can correspond to different z-scores for different sample sizes.
This is why reporting only the difference in proportions can leave out important information. A difference of 5 percentage points is more standard deviations from the model value when sample proportions have less variability. The z-score accounts for that variability; it does not change the raw difference itself.
Worked Example: The Same Difference with Two Sample Sizes
Worked Example: A Preference for Refillable Containers
Suppose 55% of households in a region regularly use refillable containers. Consider two random samples, each selected without replacement from a population of 10,000 households. One sample has 100 households and the other has 400. In both samples, 60% use refillable containers. Find the z-score for each sample proportion and compare them.
State. For both samples, \(p=0.55\) and the observed \(\hat{p}=0.60\), so each observed difference is \(0.60-0.55=0.05\). We calculate a separate standard deviation for each sample size.
Plan and check conditions. Both samples are random. For the sample of 100, \(100\leq0.10(10{,}000)=1{,}000\); for the sample of 400, \(400\leq1{,}000\). Thus, each meets the 10% condition. For \(n=100\), the expected numbers of households using and not using refillable containers are \(100(0.55)=55\) and \(100(0.45)=45\). For \(n=400\), they are \(400(0.55)=220\) and \(400(0.45)=180\). All are at least 10, so the Large Counts condition is met for both.
Do. For the sample of 100 households:
For the sample of 400 households:
As an arithmetic check, the standard deviation for \(n=400\) is half the one for \(n=100\), because the sample size is four times as large. The same difference of \(0.05\) therefore gives a z-score about twice as large: \(2.010\) compared with \(1.005\).
Conclude. The observed proportion of 0.60 is about 1.005 standard deviations above 0.55 for the sample of 100 households, and about 2.010 standard deviations above 0.55 for the sample of 400. The proportions have the same raw difference from \(p\), but the larger sample has less sampling variability.
Common Mistakes and AP Exam Communication
A complete response shows which proportion is observed, which is the model parameter, and which standard deviation belongs in the calculation. Then it translates the signed result into a statement about the sampling distribution.
- Reversing the subtraction. Use \(\hat{p}-p\). Reversing it changes the sign and therefore changes whether the result is above or below the parameter.
- Using the count instead of the proportion. The formula compares \(\hat{p}\) with \(p\). If the sample contains 47 successes out of 100, first calculate \(\hat{p}=0.47\); do not substitute 47.
- Using the wrong standard deviation. Use \(\sqrt{p(1-p)/n}\) for the sampling distribution of \(\hat{p}\). The standard deviation of individual 0-or-1 responses is not the standard deviation of the sample proportion.
- Substituting the observed proportion for \(p\). In this setting, \(p\) is the population proportion specified by the model. The formula’s standard deviation is calculated from that \(p\), not from \(\hat{p}\).
- Calling the z-score a probability or percentage-point difference. A z-score is a signed distance in standard-deviation units. State the raw difference separately if the question asks for it.
- Skipping model conditions. Show the random sampling basis, check the 10% condition when sampling without replacement, and verify \(np\geq10\) and \(n(1-p)\geq10\) before relying on a Normal model.
Key Takeaway
A z-score standardizes the difference between an observed sample proportion and the population proportion. Its sign gives the direction, and its magnitude gives the distance in standard deviations of the sampling distribution. Check the sampling conditions before treating that standardized value as a score on the standard Normal scale.
Check Your Understanding
Use the sample-proportion z-score formula and explain each result in context.
- A random sample of 120 homes is selected without replacement from a population of 4,000 homes. The model proportion is \(p=0.30\), and 42 sampled homes have a garden. Check the conditions, calculate the z-score, and interpret it.
- A random sample of 150 people is selected from a large population with \(p=0.20\). The sample proportion is \(0.16\). Find the standard deviation and z-score, and explain the sign.
- Explain why a z-score is not the same thing as the difference \(\hat{p}-p\).
- For the same \(p\) and the same difference \(\hat{p}-p\), what happens to the magnitude of the z-score when the sample size increases? Explain using the standard deviation formula.
- What condition checks are needed before using a Normal model to interpret a sample-proportion z-score?