Tutorials › AP Statistics › Checking Conditions for Each Procedure Family

Statistical practices and exam synthesis · Tutorial 1010 of 1020

Checking Conditions for Each Procedure Family

Build a practical conditions checklist for inference, then apply it to proportion, t, and chi-square scenarios using evidence from the study design and data.

Intermediate 9 min read

What You'll Learn

  • Distinguish random sampling from random assignment and explain what each supports.
  • Apply the 10% condition when sampling without replacement.
  • Check Large Counts conditions differently for proportion tests and intervals.
  • Assess normality for one-sample, paired, and two-sample t procedures.
  • Verify expected-count and independence conditions for chi-square tests.
  • Write condition checks that cite specific facts from a study or sample.

Procedure Choice Is Only the First Check

In “Decision Flowchart for Inference Scenarios,” you matched a question and data structure to an inference procedure. Before carrying out that procedure, check whether its conditions are plausible. A procedure name does not establish that the data meet the requirements for using it.

The checks usually concern how the data were collected, whether observations can reasonably be treated as independent, and whether the sample information supports the procedure’s model. The details differ across proportion \(z\), \(t\), and chi-square procedures. In particular, “large enough” means something different for Large Counts in a proportion procedure than it does for the approximate normality condition in a \(t\) procedure.

Key idea: A useful condition check names the condition and points to evidence in the context or data. “The conditions are met” is not enough: say what was random, how independence is supported, and which count or distribution check applies.

A Checklist for the Main Procedure Families

Start with the data-collection design. A random sample helps support inference about the population from which the sample was selected. Random assignment in an experiment helps support a cause-and-effect conclusion about the treatments. These are not interchangeable: randomly assigning volunteers to treatments does not, by itself, make them representative of a wider population.

Next, consider independence. If individuals are sampled without replacement, the 10% condition is a common way to justify treating observations as approximately independent: the sample size should be no more than 10% of the population size. For two samples, check this separately for each population when applicable. In an experiment, explain how the assignment and study design support independent treatment groups and independent responses; do not claim that a 10% sampling condition applies when there was no population sample.

Conditions: Use this quick map, then verify the details for the specific procedure:
  • One-proportion \(z\): random sample or appropriate random assignment; independence, including the 10% condition when sampling without replacement; and Large Counts using the null proportion for a test or the sample proportion for an interval.
  • Two-proportion \(z\): random samples or random assignment to independent groups; independence within and between groups, including the 10% condition for each population sample when applicable; and Large Counts using the pooled proportion for a test or each group’s sample proportion for an interval.
  • One-sample and paired \(t\): random sample or appropriate random assignment; independence of observations (or pairs); and a distribution check for the quantitative values or, for paired data, the differences.
  • Two-sample \(t\): random samples or random assignment to independent groups; independence within each group and between groups; and a distribution check in each group.
  • Chi-square tests: random sample(s) or appropriate random assignment; independent observations, including the 10% condition for sampling without replacement when applicable; and sufficiently large expected counts, with every expected count at least 5 under the usual AP condition.

For a one-proportion \(z\) test, the Large Counts condition is checked assuming the null hypothesis is true: \(np_0 \ge 10\) and \(n(1-p_0) \ge 10\). For a one-proportion \(z\) interval, use the observed sample proportion \(\hat{p}\): \(n\hat{p} \ge 10\) and \(n(1-\hat{p}) \ge 10\). For a two-proportion test, use the pooled proportion under the null; for an interval, check successes and failures separately in each group using that group’s sample proportion.

For \(t\) procedures, assess the shape of the quantitative data or paired differences with a graph, such as a dotplot or boxplot, and any relevant information about the population. If the sample is small, the data should be roughly symmetric with no strong skewness or outliers. Larger samples give \(t\) procedures more tolerance for departures from normality, but a severe outlier or strong skewness can still be a concern. For a two-sample \(t\) procedure, assess the distributions in both groups. For paired data, assess the distribution of the differences, not the two sets of measurements separately.

For a chi-square test, expected counts are calculated from the null model, not simply copied from the observed table. For a test of independence, the expected count in a cell is \(\dfrac{(\text{row total})(\text{column total})}{\text{grand total}}\). The test of homogeneity uses the corresponding expected count calculation for the group and category totals. Then check every cell against the expected-count condition.

Worked Checks in Context

Worked Example: Checking a One-Proportion Test

Scenario. A fictional city program takes a random sample of 140 households and asks whether each has a working carbon-monoxide alarm. In the sample, 84 households answer yes. The program asks whether there is evidence that the proportion of all city households with a working alarm differs from 0.52. Assume the city has 2,400 households.

State. Let \(p\) be the proportion of city households with a working carbon-monoxide alarm. The hypotheses are \(H_0:p=0.52\) and \(H_a:p\ne0.52\). The response is binary and there is one sample, so the procedure is a one-proportion \(z\) test.

Plan and conditions. The households were selected by random sampling, supporting inference to city households. The sample is less than 10% of the population: \(140 \le 0.10(2400)=240\), so the 10% condition is met. For the Large Counts condition, calculate using the null value:

$$ np_0=140(0.52)=72.8 \ge 10,\qquad n(1-p_0)=140(0.48)=67.2 \ge 10. $$

Both expected counts are at least 10. The conditions for the one-proportion \(z\) test are met.

Do. The sample proportion is \(\hat{p}=84/140=0.60\). The test statistic is

$$ z=\frac{\hat{p}-p_0}{\sqrt{\frac{p_0(1-p_0)}{n}}} =\frac{0.60-0.52}{\sqrt{\frac{0.52(0.48)}{140}}} \approx 1.894. $$

The two-sided \(p\)-value is approximately 0.0582, rounded to four decimal places. At the 0.05 significance level, this \(p\)-value is greater than 0.05, so we fail to reject \(H_0\).

Conclude. The sample does not provide convincing evidence that the proportion of city households with a working carbon-monoxide alarm differs from 0.52. This conclusion is about the city households represented by the random sample.

Worked Example: Checking a Two-Sample \(t\) Procedure

Scenario. In a fictional randomized experiment, 16 volunteers are assigned to use a new stretching routine and 14 different volunteers are assigned to use their usual warm-up. Their flexibility change, measured in centimeters, is recorded after four weeks. The question asks whether the mean change differs between the two routines. Dotplots of the changes show no strong skewness or outliers in either group.

Procedure. The response is quantitative, the groups contain different volunteers, and the question compares two population means. The appropriate method is a two-sample \(t\) test, not a paired \(t\) procedure.

Randomization and independence. Volunteers were randomly assigned to one of two groups, which supports a treatment comparison. Each volunteer appears in only one group, so the groups are independent by design. The description treats volunteers’ responses as independent within groups; if the same volunteer had supplied both measurements, this would instead be a paired design.

Normality check. The sample sizes are \(n_1=16\) and \(n_2=14\), both relatively small. Therefore, the distribution condition matters: the dotplots show no strong skewness or outliers in either group, supporting the use of the two-sample \(t\) procedure. It would not be enough to inspect only the combined data, because the procedure compares the two groups.

Conclusion about the plan. The random assignment, independent groups, and distribution evidence support proceeding with the two-sample \(t\) test. Random assignment supports a conclusion about the difference caused by these routines for volunteers like those in the experiment; it does not by itself establish that the volunteers represent all people.

Worked Example: Checking Expected Counts for a Chi-Square Test

Scenario. A fictional random sample of 120 library visitors is classified by visit type (first-time or returning) and whether the visitor used a study room (yes or no). The observed counts are shown below. The question is whether visit type and study-room use are associated among library visitors.

Visit typeUsed study roomDid not use study roomTotal
First-time303060
Returning204060
Total5070120

Procedure and design. One sample is classified by two categorical variables, and the question asks whether they are associated. Use a chi-square test of independence. The visitors were randomly sampled, and the observations are individual visitors, each counted in one cell. If the sample was taken without replacement, the 10% condition requires at least 1,200 visitors in the population; the scenario should provide enough population information to verify this condition rather than assuming it.

Expected-count check. Under the null hypothesis of no association, calculate each expected count from its row and column totals. For first-time visitors who used a room:

$$ E=\frac{(60)(50)}{120}=25. $$

The other expected counts are \((60)(70)/120=35\) for first-time visitors who did not use a room, \((60)(50)/120=25\) for returning visitors who used a room, and \((60)(70)/120=35\) for returning visitors who did not use a room. The expected counts are 25, 35, 25, and 35; every one is at least 5. The expected-count condition is met. The 10% condition cannot be declared met without knowing that the population of visitors is at least 1,200.

Conclusion about the plan. The random sample and expected counts support the chi-square test, provided the population-size information confirms the 10% condition when sampling without replacement. Stating that qualification is more accurate than claiming all conditions are satisfied from the information given.

Worked Example: Checking a Paired \(t\) Procedure

Scenario. Ten gardeners measure the number of minutes it takes them to prepare a raised bed before and after using a new tool. Each gardener supplies both times. A dotplot of the ten differences, calculated as after minus before, is roughly symmetric and has no apparent outliers. The gardeners were selected at random from a club of 180 members.

Procedure. The same gardeners provide both quantitative measurements, so the observations are paired. Define one difference for each gardener and use a paired \(t\) procedure for the population mean difference \(\mu_d\).

Conditions. The random sample supports inference to members of this club. The sample is below 10% of the club: \(10 \le 0.10(180)=18\). The relevant distribution check is on the ten differences, not on the before-times and after-times separately. Their dotplot is roughly symmetric with no apparent outliers, which supports using a \(t\) procedure with this small sample. These checks support the procedure for the mean change in preparation time among club members.

Common Mistakes and What Full Credit Requires

  • Saying “random” without explaining what was randomized. Identify whether the study used random sampling or random assignment. Random sampling supports population inference; random assignment supports causal inference about treatments.
  • Applying the 10% condition automatically. It is relevant to sampling without replacement from a finite population. State the population size and compare it with the sample size; do not invent a population size or use the condition as a substitute for a design explanation.
  • Using the wrong counts for a proportion procedure. For a test, use the null proportion (or pooled null proportion for a two-proportion test). For an interval, use observed sample proportions. Show the success and failure counts that establish the condition.
  • Checking the wrong distribution for a \(t\) procedure. For paired data, examine the differences. For two independent samples, assess each group. A large combined sample does not make a strong outlier in a small group irrelevant.
  • Checking observed counts instead of expected counts for chi-square. The chi-square condition concerns expected counts under the null model. Show the calculation or otherwise demonstrate that every expected count meets the threshold.
  • Claiming a condition is met when information is missing. If a scenario does not state that a sample is random or give enough information to verify the 10% condition, say so. Identify what additional information would be needed.

A strong AP response is specific and appropriately limited: “The sample was randomly selected, and \(n=140\) is less than 10% of the stated population of 2,400. Under \(H_0\), the expected success and failure counts are 72.8 and 67.2, both at least 10.” For a \(t\) procedure, cite the relevant plot and sample size; for chi-square, show that every expected cell count meets the condition.

Key takeaway: Check the design first, then independence, then the procedure-specific model condition. Use null-based counts for proportion tests, sample-based counts for proportion intervals, the correct quantitative distributions for \(t\) procedures, and expected counts for chi-square tests. Support each statement with evidence from the scenario.

Check Your Understanding

For each situation, identify the relevant conditions and state what information or evidence would let you verify them.

  1. A random sample of 90 residents is asked whether they compost. The town has 1,500 residents. What design and 10% checks apply to a one-proportion \(z\) interval?
  2. Two independent random samples are used for a two-proportion \(z\) test. What proportions should be used for the Large Counts checks under the null hypothesis?
  3. The same 12 runners record their times before and after a training plan. Which values should be examined for the \(t\) condition, and what features would be concerning?
  4. A chi-square test of independence has several observed cell counts below 5. Does that establish a condition problem? What counts should be calculated instead?
  5. A randomized experiment assigns volunteers to one of two treatments but uses no random sample from a larger population. What kind of inference does random assignment support, and what claim does it not establish by itself?