Check the Evidence, Not Just the Labels
Condition checks can look simple: identify the procedure, substitute values, and decide whether each requirement is met. But a calculation can be arithmetically correct and still answer the wrong question. In Spotting Violated Conditions in Study Descriptions, you practiced deciding whether a condition is met, violated, or not established from the study description. Here, we focus on common errors in making those decisions.
For a one-proportion \(z\)-test, two mistakes are especially tempting: treating \(n\) itself as a Large Counts check, and using \(\hat{p}\) rather than the null value \(p_0\). A third mistake is to notice the numbers but ignore how the observations were obtained. A complete check connects the proposed inference to the study’s sampling or assignment method, source population, and relevant counts.
The earlier tutorials Checking the Success-Failure Condition for Tests and Interval Versus Test Condition Checks Compared explain why the test and interval checks differ. Use those rules here as a way to audit your work: first identify the procedure, then make sure the evidence matches it.
A Three-Part Audit for Condition Checks
Before calculating, identify what the proposed procedure is intended to learn. A test evaluates a claim about a population proportion under a null hypothesis. An interval estimates a population proportion. That distinction tells you whether Large Counts is based on expected counts under \(H_0\) or observed counts in the sample.
Is the proposed analysis a one-proportion test or a one-proportion interval? Write that down before choosing counts.
For a test, use \(p_0\) from the null hypothesis. For an interval, use \(\hat{p}\), or equivalently the observed successes \(x\) and failures \(n-x\).
Explain whether the data came from random sampling or an appropriate randomized process, and connect that method to the population or process the conclusion targets.
The 10% condition has its own evidence: when sampling without replacement from a finite population of size \(N\), compare the sample size with 10% of that actual source population. It is not another Large Counts calculation, and it does not repair a problem with randomness.
| Proposed procedure | Large Counts evidence | What not to substitute |
|---|---|---|
| One-proportion \(z\)-test of \(H_0:p=p_0\) | Expected counts \(np_0\) and \(n(1-p_0)\) | \(n\) alone or counts based on \(\hat{p}\) |
| One-proportion \(z\)-interval | Observed counts \(x=n\hat{p}\) and \(n-x\) | Expected counts based on a claimed \(p_0\) |
The table concerns only Large Counts. For either procedure, also assess Random and, when appropriate, the 10% condition. Passing Large Counts does not establish that the data represent the population named in the conclusion.
Worked Examples
Worked Example: “The Sample Is Large” Is Not the Test Check
A quality manager wants to test whether more than 12% of the labels in a large shipment contain a printing error. A random sample of 75 labels is selected without replacement from a shipment of 5,000 labels. Fourteen sampled labels contain an error. Before conducting the test, a student writes, “Large Counts is satisfied because \(n=75\), which is greater than 10.” Diagnose the condition check and give a complete response about whether a one-proportion \(z\)-test is supported.
State: Let \(p\) be the proportion of labels in this shipment that contain a printing error. The hypotheses are \(H_0:p=0.12\) and \(H_a:p>0.12\).
Plan: A one-proportion \(z\)-test would require the Random condition, the 10% condition for sampling without replacement, and the Large Counts condition based on the null proportion. The statement \(n=75>10\) does not check Large Counts: that condition concerns expected successes and failures under \(H_0\), not whether the sample size exceeds 10.
Do: The labels were randomly sampled, so the Random condition is met for inference about the shipment. For the 10% condition, \(0.10(5{,}000)=500\), and \(75\leq500\), so it is met. For Large Counts, use \(p_0=0.12\):
The expected number of errors is 9, which is less than 10. The expected number without errors is 66, which is at least 10. Because both expected counts must be at least 10, Large Counts is not met. The observed count of 14 errors does not change this test condition.
Conclude: A one-proportion \(z\)-test is not supported by all the stated conditions because the null model predicts only 9 errors. The student’s conclusion based on \(n=75\) is incorrect. This condition diagnosis does not by itself establish whether the proportion exceeds 12%; it says the usual Normal-based test is not justified by the Large Counts check.
Worked Example: Using \(\hat{p}\) in a Test Check
A recycling coordinator tests whether 30% of households in a district separate food scraps for composting. A random sample of 50 households is taken without replacement from a district list of 900 households. Five sampled households separate food scraps. A student uses \(\hat{p}=5/50=0.10\), then checks \(50(0.10)=5\) and \(50(0.90)=45\), concluding that Large Counts fails for the test. Is that the right condition check?
No. The proposed analysis is a one-proportion \(z\)-test of \(H_0:p=0.30\), so the Large Counts check uses \(p_0=0.30\), not the observed \(\hat{p}=0.10\). Under the null hypothesis:
Both expected counts are at least 10, so Large Counts is met for the test. The student calculated the observed counts correctly—5 successes and \(50-5=45\) failures—but used an interval-style count check for a test.
The other conditions should still be checked. The sample was randomly selected, so Random is met for the district list. Ten percent of the source population is \(0.10(900)=90\), and \(50\leq90\), so the 10% condition is met. Thus, the stated conditions support using a one-proportion \(z\)-test. Whether the data provide convincing evidence against \(H_0\) is a separate question; passing conditions does not decide the test’s result.
Worked Example: Correct Counts for an Interval, Limited Generalization
A company invites employees to volunteer for a study of a new safety reminder. The 160 volunteers are randomly assigned in equal groups to receive either the new reminder or the usual reminder. In the new-reminder group, 60 of the 80 volunteers follow the requested safety step. A student proposes a one-proportion confidence interval for the proportion of all company employees who would follow the step with the new reminder. The student counts 60 successes and \(160-60=100\) failures. Identify the count error and assess what randomization supports.
For the new-reminder group, the relevant sample size is \(n=80\), not 160. The 160 volunteers were divided into two groups, and the outcome information given here is only for the 80 assigned the new reminder. The correct observed counts for that group are:
Both observed counts are at least 10, so the Large Counts condition for an interval is met for the new-reminder group’s observed outcomes. The calculation \(160-60=100\) is arithmetically correct but does not give the failure count for the group of 80 receiving the new reminder. The outcomes for the other 80 volunteers are not provided, so no success or failure count for that group can be calculated.
Random assignment is important, but it does not mean these volunteers were randomly sampled from all company employees. As discussed in Checking Conditions for Data from an Experiment, random assignment can support conclusions about an experimental treatment for the units in the experiment. It does not, by itself, make a volunteer group representative of all employees. Therefore, the study description does not justify generalizing an interval from these volunteers to all company employees. Also, without the other group’s outcomes, the information given cannot establish a treatment difference.
The precise diagnosis separates three points: use \(n=80\) and 20 observed failures for the new-reminder group; Large Counts passes for that group’s interval check; and volunteer recruitment limits claims about all employees. Correct counts do not remove a limitation in how participants entered the study.
Worked Example: Random Invitations Do Not Guarantee Random Responses
A library randomly selects 400 cardholders from its membership list and emails them a question about a proposed extended-hours schedule. Only 120 people respond. Among the respondents, 78 support the schedule and 42 do not. The library proposes a one-proportion confidence interval for the proportion of all cardholders who support the schedule. A student notes that the original invitations were random and checks only that \(n=120\) is large. What has the student missed?
The student has not traced who supplied the data. The people who answered chose whether to respond; the 120 respondents are not automatically a random sample just because the 400 invitations were selected at random. Nonresponse could be related to opinions about the schedule, so the Random condition for treating respondents as representative of all cardholders is not established.
For an interval, use the observed counts among the 120 respondents: \(x=78\) supporters and \(120-78=42\) who do not support the schedule. Both are at least 10, so the Large Counts condition is met. This is the correct check; merely saying \(n=120\) is “large” does not show the two observed counts.
If the 400 were sampled without replacement from a finite list of \(N\) cardholders, the 10% condition concerns the sample of 400 relative to that list: verify \(400\leq0.10N\). The respondent count of 120 is not a substitute for the original selection size in this comparison. If the list size is not given, the numerical 10% check cannot be verified from these facts. Even if it passes, it does not resolve the concern that some selected cardholders did not respond.
The interval can describe the respondents’ support rate, \(78/120=0.65\), or 65%. But the stated facts do not establish that the usual interval represents all cardholders. A careful answer distinguishes what the counts support from what the response process leaves uncertain.
Common Mistakes and What Full Credit Requires
- Using \(n\) as the Large Counts result. Saying “\(n=75\), so the condition passes” does not check expected successes and failures. For a test, show \(np_0\) and \(n(1-p_0)\), then compare each with 10.
- Using \(\hat{p}\) in a test’s count check. The sample proportion describes the observed result, but the test’s Large Counts check asks what counts the null model predicts. Use the \(p_0\) in \(H_0\).
- Using test counts for an interval—or interval counts for a test. For an interval, check \(x\) and \(n-x\). For a test, check \(np_0\) and \(n(1-p_0)\). Name the procedure so the reader can see why you chose those quantities.
- Mixing groups or denominators. In a study with treatment groups, use the sample size and outcomes for the group relevant to the proposed one-proportion calculation. Do not assign outcomes to people whose outcomes were not reported.
- Ignoring how participants entered the study. Random assignment is not random sampling, and randomly selecting people to contact is not the same as randomly selecting the people who respond. State what was randomized and limit the conclusion accordingly.
- Treating a passing numerical check as proof that inference is appropriate. Large Counts and the 10% condition do not fix voluntary response, an unsuitable sampling frame, or another problem with the Random condition. Assess each condition on its own evidence.
Key Takeaway
A reliable condition check is an evidence audit. For a test, the null proportion determines the expected counts; for an interval, the sample determines the observed counts. Then examine how the data were obtained and what population the study can represent. Keep the group, denominator, and conclusion aligned.
Check Your Understanding
For each scenario, identify the common condition-checking error, if any, and explain what evidence should be used instead.
- A random sample of 60 items is tested under \(H_0:p=0.15\). A student says Large Counts passes because \(n=60\). Calculate the correct expected counts and decide whether the condition passes.
- A random sample of 80 people has 12 successes. For a test of \(H_0:p=0.25\), a student checks \(80(12/80)\) and \(80(1-12/80)\). What should be checked for the test?
- A one-proportion interval is proposed for a group of 70 participants, of whom 55 had the characteristic. Which observed counts belong in the interval’s Large Counts check?
- Researchers randomly assign 100 volunteers to two groups, but report outcomes only for one group of 50. What sample size and outcome counts can be used for an interval about that group? What cannot be concluded about all volunteers from random assignment alone?
- A random sample of 300 cardholders is invited to respond, and 90 respond. Why does the random selection of invitees not automatically establish the Random condition for an interval based on the respondents?