A Procedure Is Only as Good as Its Conditions
In “Choosing Between Mean and Proportion Procedures,” the response and parameter guided the choice between a t procedure for a mean and a z procedure for a proportion. For mean inference, the study design then identifies whether to use one-sample t, paired t, or unpooled two-sample t. Before calculating or interpreting results, check that the chosen procedure’s conditions fit how the data were collected and what the data look like.
Some conditions concern the design. Were units selected randomly or treatments assigned randomly? Are observations independent? Other conditions concern the data’s distribution. For a small sample, is the distribution reasonably free of strong skew and outliers? These checks are not interchangeable: a graph cannot establish that sampling was random, and random selection does not ensure that a small sample has a suitable shape.
Check the Design: Randomness and Independence
For a t procedure to support inference, the data should come from an appropriate random sample or randomized experiment. A random sample helps justify conclusions about the population represented by the sampling method. Random assignment helps justify a comparison of treatments in an experiment. They serve different purposes: random assignment alone does not make the experimental units a random sample of a broader population.
Also check that observations are independent. One unit’s measurement should not mechanically determine another unit’s measurement. Repeated measurements on the same person, or deliberately matched subjects, are linked; they are not independent observations for a two-sample t procedure. As in “Spotting Paired Designs in Word Problems,” analyze genuine pairs through their differences. For paired t, the differences from distinct pairs should be independent.
When units are sampled without replacement from a finite population, use the 10% condition: the sample size should be no more than 10% of the population size. For two independent random samples, check this condition separately for each population when applicable. In a randomized experiment, the 10% condition is not a substitute for checking whether the assignment and observations support the analysis.
- Identify the random process: random sampling, random assignment, or both.
- Check that the observations used as separate data points are independent under the study design.
- For sampling without replacement, check the 10% condition against the relevant population size.
- For paired t, check independence between pairs, not between the two measurements within a pair.
Check the Shape: Which Data Need to Look Reasonable?
A t procedure uses a t distribution to account for the fact that the population standard deviation is generally unknown and estimated with the sample standard deviation \(s\). Its reliability also depends on the sample size and the shape of the data. For a small sample, inspect a dotplot, stemplot, or other suitable graph for strong skewness or outliers. If the sample is sufficiently large, the sampling distribution of the sample mean is generally more nearly normal, but a very severe outlier or extreme skew can still be concerning.
For a one-sample t procedure, examine the distribution of the individual measurements in the sample. For paired t, the procedure is applied to the differences, so examine the distribution of those differences—not the before measurements and after measurements separately. For two-sample t, examine the response distribution within each group. A combined graph can hide a problem in one group.
A practical AP Statistics check for a small sample is whether the relevant graph looks roughly symmetric or otherwise reasonably well behaved, with no strong outliers. A sample size around 30 or more in the relevant group often provides some protection against moderate departures from normality. It is not a magic guarantee: consider the actual shape, especially when it is strongly skewed or contains an extreme outlier. State what the graph shows rather than claiming that a small sample is “normal” without evidence.
- One-sample t: assess the sample’s individual quantitative measurements.
- Paired t: assess the pairwise differences, defined in a stated order.
- Two-sample t: assess each group’s measurements separately.
- With a small sample, look for a reasonably well-behaved shape and no strong outliers. With larger samples, moderate non-normality is less concerning, but severe skewness and extreme outliers still deserve attention.
If a condition is not supported, do not simply proceed as if it were met. Explain which detail is concerning and why it matters. A different procedure is not automatically justified; the next step depends on the study question and design.
Worked Examples: Match Each Check to the Study
Worked Example: A One-Sample t Test for Garden Yields
A fictional agricultural program wants to know whether the mean yield per garden bed exceeds 68 kilograms. Staff randomly select 36 beds from 720 beds in the program. The sample mean yield is 72 kilograms, and the sample standard deviation is 12 kilograms. A dotplot is roughly symmetric, with no apparent outliers. Test the claim at \(\alpha=0.05\).
State: Let \(\mu\) be the true mean yield, in kilograms, per bed in the program. The hypotheses are \(H_0:\mu=68\) and \(H_a:\mu>68\).
Plan: Use a one-sample t test because one quantitative sample is being compared with a fixed value. The beds were randomly selected. The 10% condition holds because \(36\leq0.10(720)=72\). The dotplot is roughly symmetric with no apparent outliers, so the shape is reasonable for a t procedure. The procedure is appropriate based on the stated design and graph.
Do: The standard error is \(s/\sqrt{n}=12/\sqrt{36}=12/6=2\) kilograms. The test statistic is
With \(df=36-1=35\), the right-tailed p-value is approximately \(0.0267\), rounded. A calculator can obtain it with a t-distribution area to the right of 2.00 with 35 degrees of freedom.
Conclude: Since \(0.0267<0.05\), reject \(H_0\). The data provide convincing evidence that the mean yield per bed in the program exceeds 68 kilograms.
The procedure choice depends on more than the arithmetic: the random sample supports inference about the program’s beds, the 10% condition supports independence when sampling without replacement, and the graph supports using a t procedure.
Worked Example: Check the Differences in a Paired Study
A fictional clinic records the time, in minutes, that 14 volunteers need to complete a task before and after a brief training session. The same volunteers are measured twice. The research question concerns whether the mean time changes. A graph of the differences, defined as after minus before, shows a strong right skew and one unusually large positive difference.
Identify the design: The two measurements come from the same volunteer, so they are paired. Define \(d=\text{time after}-\text{time before}\); the target is the population mean difference \(\mu_d\). If the conditions were supported, the appropriate mean procedure would be paired t applied to the 14 differences, not two-sample t applied to the before and after measurements as independent groups.
Check the design conditions: The volunteers were recruited by an open invitation, not selected randomly. The sample therefore does not support generalizing to all clinic volunteers as though they were randomly sampled. The measurements within a volunteer are intentionally linked; independence is relevant between distinct volunteers or pairs. The study does not state that the volunteers were randomly assigned to a treatment, so the design also does not establish a randomized experiment.
Check the shape condition: There are only 14 differences, and their graph shows strong skew and an unusually large value. These features make the t procedure’s shape condition questionable. The before and after measurements might each have a different appearance; that would not resolve the issue because paired t uses the distribution of the differences.
Conclusion about the procedure: Paired t is the procedure that matches the linked quantitative response and target, but the stated recruitment and difference plot do not support a routine paired t analysis for broad population inference. Report these limitations rather than claiming that the procedure’s conditions have been met. More information about the sampling process or additional data could change the assessment.
Worked Example: Two Independent Treatment Groups
In a fictional experiment, researchers randomly assign 64 seedlings to two watering schedules, with 32 seedlings in each group. After six weeks, they measure plant height in centimeters. Researchers want to compare the population mean heights under the two schedules. Within each group, the height distributions are mildly right-skewed, with no apparent outliers.
Identify the procedure: Height is quantitative, and different seedlings receive the two schedules. There are no described pairs, so the target is the difference between the two population mean heights, \(\mu_1-\mu_2\). Use an unpooled two-sample t procedure, as in “Experiments Comparing Two Treatments.”
Check the design: Seedlings were randomly assigned to treatments, which supports comparing the treatments in this experiment. Each seedling contributes one measurement to one group, so the groups are independent by design. The study describes random assignment, not random sampling from all seedlings; conclusions should not be generalized to an entire broader population solely on the basis of this experiment.
Check the data shape: Examine the height distribution in each group. Both groups have 32 observations, mild right skew, and no apparent outliers. These sample sizes and shapes do not reveal a serious concern for using the t procedure. Mild skew is not the same as extreme skew, and no group has an identified outlier that would dominate its mean.
Conclusion about the procedure: The unpooled two-sample t procedure matches the independent treatment groups and quantitative response, and the stated assignment and graph support its use. The random assignment supports a treatment comparison, while the lack of random sampling limits generalization beyond the experimental units represented.
Common Mistakes and AP Exam Tips
- Checking the wrong graph: For paired t, inspect the differences. For two-sample t, inspect each group. A graph of all measurements together can conceal the relevant shape.
- Treating random assignment as random sampling: Random assignment supports a treatment comparison; random sampling supports generalizing to a population represented by the sample. Name the process the study actually used.
- Applying the 10% condition to every design without explanation: It applies when sampling without replacement from a finite population. State the sample size and population size when checking it.
- Assuming a large sample fixes every problem: Larger samples help with moderate non-normality, but an extreme outlier or severe skew still calls for attention. Describe the evidence rather than relying on a cutoff alone.
- Calling observations independent because they are in separate columns: Before-and-after measurements on the same person are linked. The data layout does not override the study design.
- Writing only “conditions are met”: Full-credit communication identifies each condition and the study detail that supports it, such as “36 of 720 beds were randomly selected, so \(36\leq72\) for the 10% condition.”
A careful plan names the procedure and connects each check to evidence: how units were sampled or assigned, why observations are independent, whether the 10% condition applies, and what the relevant graph shows. If a condition is questionable, state the concern precisely and limit the conclusion accordingly.
Check Your Understanding
For each situation, identify the condition evidence you would check before using the proposed mean procedure.
- A random sample of 24 residents is selected without replacement from a town of 500. What does the 10% condition require, and which graph feature matters for one-sample t?
- The same 12 runners record their times before and after a training plan. Which observations should a paired t shape check examine?
- Two independent random samples are taken from populations of 200 and 800. The sample sizes are 18 and 50. Which population sizes should be used for the 10% checks?
- A randomized experiment assigns different students to two study methods. Explain what random assignment supports and what it does not establish about generalizing to all students.
- Each of two independent groups has 35 observations, but one group’s graph shows an extreme outlier. Why is “both sample sizes exceed 30” not a complete condition check?