Put Procedure Choice and Regression Reasoning Together
Mixed practice asks you to shift gears. One part may ask whether a population proportion differs from a claimed value; another may ask which procedure fits a study design; a later part may ask what a correlation or regression equation says about an association. The challenge is not to use every method you know. It is to match each question to the evidence and method it calls for.
In “Choosing the Correct Inference Procedure” and “Writing a Complete Inference Response,” you learned to identify the response type and data structure, then organize inference as State, Plan, Do, and Conclude. In the regression tutorials, you practiced interpreting \(r\), \(r^2\), the slope, the intercept, and \(s\). This timed set brings those skills together. It does not ask for inference about a regression slope.
A Timed Routine for a Mixed Set
Try the three prompts below in about 20 minutes: allow roughly 6 minutes for each prompt and keep 2 minutes for a final check. The timing is a guide, not a reason to skip statistical reasoning. If one calculation stalls, record the procedure or formula you know is appropriate, move to the next part, and return if time allows. This builds on “Time Management on the Digital Exam” and “Approaching a Multi-Part Free-Response Question.”
Identify whether it asks about a proportion, a mean, a relationship between categorical variables, or a quantitative association.
For inference, distinguish one sample, two independent groups, and paired observations. For regression, identify the explanatory and response variables.
Name the procedure or statistic, check relevant conditions, and show the calculation or output needed to answer the question.
Interpret the result in the situation, with the relevant population or units, and keep the claim within the study’s design and scope.
Worked Set
Worked Example: Test a Claim About a Population Proportion
Prompt. A manufacturer claims that 55% of its battery packs pass a quality check. In a hypothetical audit, a random sample of 120 packs from a production run of 1,800 includes 76 that pass. Is there convincing evidence that the proportion of all packs in this run that pass is greater than 55%? Use a significance level of 0.05.
State. Let \(p\) be the proportion of all 1,800 battery packs in this production run that pass the quality check. The hypotheses are \(H_0:p=0.55\) and \(H_a:p>0.55\). The alternative matches the question’s claim of a greater pass proportion.
Plan. Use a one-proportion \(z\) test. The audit is described as a random sample. The 10% condition is met because \(120\leq0.10(1800)=180\), supporting independence when sampling without replacement. Under the null hypothesis, the Large Counts condition is met: \(np_0=120(0.55)=66\geq10\) and \(n(1-p_0)=120(0.45)=54\geq10\).
Do. The sample proportion is \(\hat{p}=76/120\approx0.6333\). Using the null value in the standard error gives:
The one-sided p-value is \(P(Z\geq1.835)\approx0.0333\), rounded to four decimal places. Since \(0.0333<0.05\), reject \(H_0\).
Conclude. The audit provides convincing evidence that more than 55% of the battery packs in this production run pass the quality check. This conclusion concerns the run represented by the random sample; it does not establish the pass proportion for every future production run.
Worked Example: Choose and Carry Out a Paired t Test
Prompt. A horticulture student randomly selects 10 garden plots from a collection of 150 plots and measures the height of the same seedlings before and after a new watering schedule. The observed differences, defined as after minus before in centimeters, are \(0,1,2,2,2,3,3,3,4,4\). A dotplot of the differences shows no strong skew or outliers. Is there evidence that the schedule increases mean seedling height? Use \(\alpha=0.05\).
State. Let \(\mu_d\) be the population mean change in seedling height, in centimeters, for plots like those sampled. Because differences are defined as after minus before, an increase corresponds to a positive difference. Use \(H_0:\mu_d=0\) and \(H_a:\mu_d>0\).
Plan. The same plots are measured twice, so the observations are paired. Analyze one difference per plot with a paired \(t\) test; a two-sample \(t\) test would incorrectly treat the before and after measurements as independent groups. The plots were randomly sampled. The 10% condition is met because \(10\leq0.10(150)=15\). The dotplot is reported to show no strong skew or outliers, supporting use of a \(t\) procedure with this small sample.
Do. The differences sum to 24, so \(\bar{x}_d=24/10=2.4\) cm. The sum of their squares is 72. The sample standard deviation is:
With \(df=10-1=9\), the test statistic and one-sided p-value are:
The p-value is rounded to four decimal places. Since it is below 0.05, reject \(H_0\).
Conclude. The sample provides convincing evidence that the watering schedule increases mean seedling height for plots like those sampled. The result supports a conclusion about the mean change in this population of plots. Because the prompt describes a random sample, it does not by itself establish that the schedule caused the increase; that would require an appropriate treatment assignment design.
Worked Example: Interpret Correlation and a Least-Squares Line
Prompt. In a hypothetical study of four garden plots, \(x\) is the average amount of water applied each week, in liters, and \(y\) is the seedling height, in centimeters. A regression summary gives \(\bar{x}=4\), \(\bar{y}=12\), \(S_{xx}=10\), \(S_{yy}=78\), and \(S_{xy}=27\). The observed \(x\)-values range from 2 to 6 liters per week. Find \(r\) and the least-squares regression line. Interpret \(r\), \(r^2\), the slope, and \(s\), given that the regression output reports \(s\approx1.60\) cm.
Calculate \(r\) and the line. The correlation is:
The positive value indicates a strong positive linear association among these four observed plots. The slope is \(b=S_{xy}/S_{xx}=27/10=2.7\). The intercept is \(a=\bar{y}-b\bar{x}=12-(2.7)(4)=1.2\). Thus the least-squares regression line is:
Interpret the summaries. Squaring the correlation gives \(r^2=(0.9668)^2\approx0.9347\); using the unrounded correlation gives \(r^2=729/780\approx0.9346\), or about 93.46%. In this sample, about 93.46% of the variation in seedling heights is accounted for by the linear relationship with weekly water applied. The slope means that for these plots, each additional liter of water per week is associated with a predicted increase of 2.7 centimeters in seedling height. It does not mean every plant gains exactly 2.7 centimeters.
The intercept of 1.2 centimeters is the model’s predicted height at zero liters per week. Since zero is outside the observed range of 2 to 6 liters per week, this intercept has little practical meaning here. The residual standard deviation \(s\approx1.60\) cm describes the typical vertical distance of observed heights from the fitted line, in centimeters.
Keep the conclusion within scope. These summaries describe four observed plots. The strong positive correlation does not alone show that applying more water caused greater height, and a prediction should not be extended beyond the observed water range without caution. As emphasized in “Communicating a Prediction and Its Accuracy,” a fitted value is an estimate, not a guaranteed outcome.
Fast Procedure-Choice Checks
A mixed set may ask for the procedure without asking you to complete the test. Make the matching explicit, and give a brief reason based on the response and design. The following contrasts help prevent common mix-ups:
- If one random sample gives a binary outcome and the question concerns one population proportion, consider a one-proportion \(z\) procedure.
- If two independent groups each give a binary outcome and the question compares population proportions, consider a two-proportion \(z\) procedure.
- If the same individuals are measured twice, or individuals are deliberately matched, calculate one difference per pair and consider a paired \(t\) procedure for a mean difference.
- If two separate groups provide quantitative responses and the question compares population means, consider a two-sample \(t\) procedure.
- If one sample is classified by two categorical variables and the question asks whether the variables are associated, consider a chi-square test of independence.
- If the question describes an association between two quantitative variables, use regression summaries such as \(r\), \(r^2\), and the least-squares line as appropriate; do not substitute an inference procedure for a descriptive question.
These are starting points, not permission to skip conditions. As in “Checking Conditions for Each Procedure Family,” verify the randomization or sampling process, independence and the 10% condition where relevant, and the appropriate distribution condition for the chosen procedure. The study description may not provide enough information to verify a condition; say so rather than assuming it is true.
Common Mistakes and AP Exam Tips
- Choosing by topic words alone. “Before and after” signals paired data when the same units are measured twice. Identify the units and how observations relate before selecting a procedure.
- Using the sample proportion in the null standard error. For a one-proportion \(z\) test, the standard error is calculated using the null proportion \(p_0\). Show that substitution so the test statistic is checkable.
- Listing conditions without checking them. A full-credit Plan names the condition and uses the scenario to verify it, such as \(120\leq0.10(1800)\), rather than merely writing “10% condition.”
- Reporting a p-value as the probability that the null is true. A p-value is the probability, assuming the null hypothesis is true, of obtaining a result at least as extreme as the one observed in the direction specified by the alternative.
- Confusing \(r\) and \(r^2\). Correlation describes direction and strength of a linear association. \(r^2\) describes the proportion of response variation accounted for by the linear model in the observed data.
- Giving a slope without units or context. State the predicted response change for each one-unit increase in the explanatory variable, using both variables’ context and units.
- Writing more than the evidence supports. A random sample can support generalizing to its population when the plan is carried out properly; random assignment can support a causal conclusion. Neither is established merely by a small p-value or a large \(r\).
- Rushing past the final sentence. A calculation without a conclusion leaves the question unanswered. State whether the evidence is convincing, identify the population parameter or observed association, and keep the claim appropriately limited.
For full credit under time pressure, make the structure visible. In an inference response, write the parameter and hypotheses, name the procedure, check each relevant condition, show the statistic and p-value, and conclude in context. In a regression response, identify which summary answers which request and include units and scope in the interpretation.
Check Your Understanding
For each question, identify the procedure or statistic needed and state the key reasoning before calculating.
- A random sample of 90 repaired phones includes 63 that work after one week. What inference procedure would test whether the population proportion working after one week differs from 0.70? Name one condition that must be checked.
- The same 18 runners record their recovery times before and after a training plan. The question asks whether mean recovery time changed. Why is a paired procedure appropriate, and what quantity should be analyzed?
- Two independent classes report whether they completed an optional review session. The question compares the completion proportions. Which procedure family fits, and what sampling or assignment details would you check?
- A regression line predicting plant height from weekly water use has slope 1.8 centimeters per liter. Write an interpretation of the slope that includes the variables and units.
- A fitted line has \(r=-0.80\). Interpret the direction and strength of the linear association, and calculate \(r^2\). What does \(r^2\) describe in context?