Tutorials › AP Statistics › Finding Percentiles of a Sample Proportion with invNorm

Sampling distributions for proportions · Tutorial 412 of 1000

Finding Percentiles of a Sample Proportion with invNorm

Use the 90th-percentile area in invNorm to find and interpret the sample-proportion value marking the top 10% of repeated samples.

Intermediate 9 min read

What You'll Learn

  • Translate “top 10%” into a cumulative area of 0.90 to the left.
  • Use invNorm with the sampling-distribution mean and standard deviation for \(\hat{p}\).
  • Check randomness, the 10% condition, and the Large Counts condition before using a Normal model.
  • Interpret the resulting cutoff as an approximate sample-proportion threshold in context.
  • Explain why an approximate cutoff need not leave exactly 10% of actual samples above it.

From a Probability to a Cutoff

In Using normalcdf for Probabilities About p-hat, you used an approximate Normal model to find the probability of sample proportions in a specified region. Now reverse the question: instead of starting with a cutoff and finding the area, start with an area and find the cutoff. The calculator function for this is invNorm.

This tutorial focuses on the value of \(\hat{p}\) that marks the top 10% of a sampling distribution. “Top 10%” means 10% of the model’s area lies to the right of the cutoff. Since invNorm takes the area to the left, its area input must be \(0.90\), not \(0.10\).

Definition: A percentile cutoff for a sampling distribution of \(\hat{p}\) is a value of \(\hat{p}\) with a specified proportion of the model’s area at or below it. The cutoff for the top 10% is the 90th percentile: 90% of the model’s area is to its left and 10% is to its right.

As in Normal Model for Sample Proportion Probabilities, when the Normal approximation is supported, the model for \(\hat{p}\) has mean \(p\) and standard deviation \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\). The new step is to give those two model parameters and the left-tail area to invNorm.

Formula: For the top 10% cutoff in an appropriate Normal model of \(\hat{p}\), use \(\operatorname{invNorm}(0.90,p,\sigma_{\hat{p}})\), where \(\sigma_{\hat{p}}=\sqrt{p(1-p)/n}\). On a TI-84, the inputs are area to the left, mean, and standard deviation.

A Setup That Keeps the Tail Straight

First translate the wording into an area statement. If \(c\) marks the top 10%, then the model assigns about \(0.10\) to values above \(c\). Equivalently, it assigns about \(1-0.10=0.90\) to values at or below \(c\). That left-tail area is what invNorm needs.

Next identify the sampling distribution, not the distribution of individual responses. For \(\hat{p}\), use the population proportion \(p\) as the mean and the standard deviation formula for sample proportions. Check the conditions for a Normal model before calculating a cutoff. The relevant checks, established in earlier tutorials, are randomness, the 10% condition when sampling without replacement, and the Large Counts condition.

1
Define the random variable.
State what \(\hat{p}\) represents for repeated random samples of the stated size.
2
Translate “top 10%.”
Use area \(0.90\) to the left, because the remaining \(0.10\) is to the right.
3
Check the Normal model conditions.
Address randomness, the 10% condition if sampling without replacement, and both expected counts.
4
Calculate the model’s parameters.
Use mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\).
5
Use invNorm and interpret.
Enter the left-tail area, mean, and standard deviation, then describe the resulting cutoff in context.

The cutoff is on the \(\hat{p}\) scale, so it is a proportion. For example, a cutoff of \(0.37\) means 37% of the sampled individuals have the characteristic at that boundary. It is not a count, and it is not the population proportion \(p\) itself.

Worked Example: The Top 10% of Sampled Gardeners

Worked Example: A Top-Ten-Percent Cutoff for Gardeners

Suppose 32% of residents in a town grow vegetables at home. A random sample of 160 residents is selected without replacement from the town’s 8,000 residents. Let \(\hat{p}\) be the proportion in the sample who grow vegetables. Find the cutoff that marks the top 10% of the sampling distribution of \(\hat{p}\).

State. The random variable \(\hat{p}\) is the proportion of residents who grow vegetables in a random sample of 160. The parameter is \(p=0.32\), and the requested value is the 90th percentile of the sampling distribution.

Plan and check conditions. The sample is random. Since it is drawn without replacement, check the 10% condition: \(160\leq0.10(8{,}000)=800\), so it is met. The expected number of residents who grow vegetables is \(np=160(0.32)=51.2\), and the expected number who do not is \(n(1-p)=160(0.68)=108.8\). Both are at least 10, so the Large Counts condition is met. The Normal approximation is supported.

Do. The mean of the sampling distribution is \(p=0.32\). Its standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{0.32(0.68)}{160}} =\sqrt{0.00136} \approx0.03687818 $$

The top 10% leaves 90% of the area to the left, so use \(0.90\) as the invNorm area:

$$ \operatorname{invNorm}(0.90,0.32,0.03687818) \approx0.3673 $$

Conclude. Under the approximate Normal model, a sample proportion of about \(0.3673\), or 36.73%, marks the top 10% of sample proportions from random samples of 160 residents. The model assigns about 10% of sample proportions to values above this cutoff.

Why the Input Is 0.90, Not 0.10

The most important decision in this calculation is choosing the area input. invNorm works from the left side of the Normal curve: it finds the value whose area to the left equals the input. The phrase “top 10%” describes the right side, so convert it to the area on the left: \(1-0.10=0.90\).

This is the same tail conversion used in Finding a Value from a Percentile with invNorm. A stated percentile is already a left-tail area when written as a proportion. A stated right-tail percentage must first be converted to its complementary left-tail area. If you enter \(0.10\) for a top-10% cutoff, you instead find the 10th percentile, a value with 10% of the model to its left.

Keep the three inputs in the calculator’s stated order: area, mean, standard deviation. The area is dimensionless; the mean and standard deviation are both measured on the \(\hat{p}\) scale. Do not enter the sample size in place of the standard deviation, and do not use the standard deviation of individual 0-or-1 responses.

Worked Example: A Small Population Proportion

Worked Example: A Top-Ten-Percent Cutoff for a Rare Preference

In a large community, 12% of residents prefer a particular type of public transit pass. A random sample of 150 residents is selected without replacement from a population of 5,000. Let \(\hat{p}\) be the sample proportion who prefer the pass. Find the cutoff that marks the top 10% of the sampling distribution.

State. Here, \(\hat{p}\) measures the proportion of the 150 sampled residents who prefer the pass. The model parameter is \(p=0.12\), and we need the 90th-percentile cutoff.

Plan and check conditions. The sample is random. For sampling without replacement, \(150\leq0.10(5{,}000)=500\), so the 10% condition is met. The expected number of residents who prefer the pass is \(np=150(0.12)=18\). The expected number who do not is \(n(1-p)=150(0.88)=132\). Both counts are at least 10, meeting the Large Counts condition.

Do. The mean is \(0.12\). The standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.12(0.88)}{150}} =\sqrt{0.000704} \approx0.02653300 $$

Use the 90% area to the left:

$$ \operatorname{invNorm}(0.90,0.12,0.02653300) \approx0.1540 $$

Conclude. The approximate cutoff is \(0.1540\), or 15.40%. Under the Normal model, about 10% of random samples of 150 residents would have a sample proportion above 15.40% who prefer this transit pass.

What the Cutoff Does—and Does Not—Guarantee

A percentile cutoff describes an area in a model, not a promise about the outcome of a particular sample. The sample proportion can only take certain values: for a sample of size \(n\), it changes in steps of \(1/n\). The Normal model is continuous and assigns area smoothly across the number line, while the actual sampling distribution of \(\hat{p}\) is discrete.

Consequently, the model-based cutoff may not be one of the possible sample proportions. Even when it is attainable, the proportion of actual samples above it need not be exactly 10%. The Normal approximation and rounding both contribute to that difference. Report the result as an approximate cutoff and avoid claiming that exactly 10% of real samples must exceed it.

A cutoff also does not change the population proportion. In the transit-pass example, \(p=0.12\) describes the population; \(0.1540\) is a threshold for sample proportions from repeated samples. The cutoff is higher than the population proportion because it is in the upper part of the sampling distribution.

Worked Example: A Cutoff Near One-Half

Worked Example: A Top-Ten-Percent Cutoff for Reusable Bottles

Suppose 47% of households in a region regularly use a reusable water bottle. A random sample of 400 households is selected without replacement from a population of 12,000 households. Let \(\hat{p}\) be the proportion of sampled households that regularly use a reusable bottle. Find the cutoff marking the top 10%.

State. The parameter is \(p=0.47\), the sample size is \(n=400\), and the random variable \(\hat{p}\) is the sample proportion of households that regularly use a reusable bottle.

Plan and check conditions. The sample is random. The 10% condition is met because \(400\leq0.10(12{,}000)=1{,}200\). The expected number of households that regularly use a reusable bottle is \(np=400(0.47)=188\); the expected number that do not is \(n(1-p)=400(0.53)=212\). Both are at least 10, so the Large Counts condition supports a Normal model.

Do. The mean is \(0.47\), and the standard deviation is:

$$ \sigma_{\hat{p}} =\sqrt{\frac{0.47(0.53)}{400}} =\sqrt{0.00062275} \approx0.02495496 $$

The top 10% cutoff has \(0.90\) of the Normal-model area to its left:

$$ \operatorname{invNorm}(0.90,0.47,0.02495496) \approx0.5020 $$

Conclude. The cutoff is about \(0.5020\), or 50.20%. Under the approximate Normal model, about 10% of samples of 400 households would have a reusable-bottle-use proportion above 50.20%.

Common Mistakes and AP Exam Communication

A strong response makes the direction of the tail and the sampling model visible. An invNorm answer by itself does not show whether the correct percentile or standard deviation was used.

  • Entering 0.10 for the top 10%. That gives the 10th percentile. The top 10% cutoff is the 90th percentile, so enter \(0.90\).
  • Using normalcdf instead of invNorm. normalcdf finds an area from a cutoff. Here the area is known and the cutoff is unknown, so use invNorm.
  • Using the wrong mean or spread. For \(\hat{p}\), use mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\). Do not substitute \(\hat{p}\), the observed sample proportion, for \(p\) in this sampling-distribution model.
  • Skipping a condition check. State why the sample can be treated as random, check the 10% condition when sampling without replacement, and verify both expected counts \(np\) and \(n(1-p)\).
  • Interpreting the cutoff as an individual-level result. The cutoff describes sample proportions from repeated samples, not the chance that one randomly selected person has the characteristic.
  • Claiming the exact top 10% of possible samples. The cutoff is based on an approximate, continuous Normal model. The actual sampling distribution is discrete, so describe the result as approximate.
AP Exam Tip: State that the requested cutoff is the 90th percentile, check the Normal-model conditions, show \(p\) and \(\sigma_{\hat{p}}\), and write the invNorm inputs in order. Conclude with the approximate sample-proportion cutoff and explain what lies above it in context.

Key Takeaway

To find the sample-proportion value marking the top 10%, use the 90th-percentile area in invNorm. First confirm that the Normal approximation for \(\hat{p}\) is supported; then use mean \(p\) and standard deviation \(\sqrt{p(1-p)/n}\). Interpret the result as an approximate threshold for sample proportions in repeated samples.

Key takeaway: “Top 10%” means \(0.90\) to the left. Use \(\operatorname{invNorm}(0.90,p,\sqrt{p(1-p)/n})\), then describe the result as an approximate cutoff for \(\hat{p}\) in context.

Check Your Understanding

For each setting, identify the left-tail area, check the Normal-model conditions, and set up or interpret the invNorm cutoff as requested.

  1. A random sample of 180 households is selected without replacement from a population of 9,000. Suppose \(p=0.35\). Check the 10% and Large Counts conditions, then write the invNorm entry for the top 10% cutoff.
  2. For a top 10% cutoff, explain why the area input is \(0.90\), not \(0.10\).
  3. A random sample of 120 people comes from a large population where \(p=0.08\). Check whether the Large Counts condition supports a Normal model for \(\hat{p}\). Should you proceed with invNorm using this model?
  4. A calculator gives a top-10% cutoff of \(0.43\). Explain what this value means for the sample proportions in repeated samples under the model.
  5. Why might the actual proportion of samples above a Normal-model top-10% cutoff differ slightly from 10%?