A Repeatable Structure for Inference Answers
In “Checking Conditions for Each Procedure Family,” you practiced verifying whether an inference method is appropriate. A complete inference response puts that work together with the calculation and a conclusion that answers the question. The goal is not to write more words for their own sake; it is to make every part of the reasoning visible and connected to the context.
For a test, a strong response usually names the parameter and hypotheses, identifies the procedure and checks its conditions, calculates the test statistic and p-value, and concludes in context. For a confidence interval, name the parameter and interval procedure, verify conditions, calculate the interval, and interpret it in context. These answers can be organized with the familiar four headings State, Plan, Do, Conclude.
The State part identifies the population parameter, not just the sample statistic. For a proportion, define \(p\); for a population mean, define \(\mu\); for a paired analysis, define the population mean difference \(\mu_d\). For a test, state the null and alternative hypotheses using that parameter. For an interval, state the parameter the interval estimates.
The Plan part names the procedure and checks its conditions with evidence. As in “Checking Conditions for Each Procedure Family,” refer specifically to the data-collection design, independence and the 10% condition when applicable, and the procedure-specific distribution or count condition. “The conditions are met” is not evidence by itself.
The Do part shows the calculation, including the formula, the values substituted, and the resulting statistic and p-value or interval. Use consistent rounding, and make clear which tail or confidence level applies. The Conclude part answers the question in context: for a test, describe the evidence about the stated claim; for an interval, interpret the range as plausible values for the population parameter.
Worked Test Response: One Population Proportion
Worked Example: Is the Citywide Proportion Different From 0.50?
Scenario. A fictional city randomly selects 100 households from its 2,200 households and asks whether each has a working smoke alarm. Sixty households answer yes. The question is whether there is evidence that the proportion of all city households with a working smoke alarm differs from 0.50. Use a significance level of 0.05.
State. Let \(p\) be the proportion of all households in this city that have a working smoke alarm. The hypotheses are \(H_0:p=0.50\) and \(H_a:p\ne0.50\). Because the question concerns one population proportion, the appropriate procedure is a one-proportion \(z\) test.
Plan. The households were randomly selected, supporting inference to the city’s households. The sample is less than 10% of the population: \(100 \le 0.10(2200)=220\), so the 10% condition is met. Under the null hypothesis, the Large Counts condition is satisfied:
The random sample, 10% condition, and null-based success and failure counts support using the one-proportion \(z\) test.
Do. The sample proportion is \(\hat{p}=60/100=0.60\). The test statistic is
For the two-sided alternative, the p-value is \(2P(Z\ge2.00)\approx0.0455\), rounded to four decimal places. Since \(0.0455<0.05\), we reject \(H_0\).
Conclude. The random sample provides convincing evidence that the proportion of city households with a working smoke alarm differs from 0.50. The sample proportion is higher than 0.50, but the test’s conclusion is that the population proportion differs; it does not establish the exact population proportion.
Notice how the conclusion follows the alternative hypothesis. Because this was a two-sided test, the conclusion addresses whether the proportion differs from 0.50. The sample result points upward, but changing the conclusion to claim the proportion is greater would require a one-sided alternative stated before the analysis.
Worked Interval Response: One Population Mean
Worked Example: Estimating the Mean Bus Commute Time
Scenario. A fictional transit agency randomly selects 25 riders from a population of 1,200 monthly riders and records each rider’s one-way commute time in minutes. The sample mean is \(\bar{x}=18.4\) minutes and the sample standard deviation is \(s=4.0\) minutes. A plot of the sample times shows no strong skewness or outliers. Construct and interpret a 95% confidence interval for the population mean commute time.
State. Let \(\mu\) be the true mean one-way commute time, in minutes, for the agency’s monthly riders. Use a one-sample \(t\) interval for \(\mu\).
Plan. The riders were randomly selected, so the sample supports inference to the agency’s monthly riders. The sample is less than 10% of the population: \(25 \le 0.10(1200)=120\), satisfying the 10% condition. The plot shows no strong skewness or outliers, supporting the distribution condition for this \(t\) procedure. With \(n=25\), checking the shape is important. These conditions support constructing the interval.
Do. The degrees of freedom are \(25-1=24\). For a 95% confidence level, the critical value is \(t^*=2.064\), rounded to three decimal places. The standard error is \(s/\sqrt{n}=4.0/\sqrt{25}=0.8\) minutes. Therefore:
The interval is \((16.7488,\ 20.0512)\) minutes, or approximately \((16.75,\ 20.05)\) minutes.
Conclude. We are 95% confident that the true mean one-way commute time for the agency’s monthly riders is between 16.75 and 20.05 minutes. The interval estimates a population mean; it does not say that 95% of individual riders have commute times in this range.
In an interval response, the confidence level describes the long-run success rate of the method: if we repeatedly took random samples in the same way and built intervals using this method, about 95% of those intervals would capture the true population mean. The particular interval either captures \(\mu\) or it does not. The standard AP interpretation is “We are 95% confident that the true mean…” followed by the parameter, population, and interval in context.
Worked Test Response: Comparing Two Proportions
Worked Example: Comparing Two Reminder Messages
Scenario. In a fictional study, independent random samples of 100 customers are drawn from each of two customer populations. One sample receives a text reminder and the other receives an email reminder. Within one week, 60 customers in the text group and 40 customers in the email group make a scheduled appointment. Assume each population contains at least 1,000 customers. Is there evidence that the appointment proportions differ between the populations? Use \(\alpha=0.05\).
State. Let \(p_1\) be the proportion of customers in the first population who would make an appointment after a text reminder, and let \(p_2\) be the corresponding proportion in the second population after an email reminder. The hypotheses are \(H_0:p_1-p_2=0\) and \(H_a:p_1-p_2\ne0\). Use a two-proportion \(z\) test.
Plan. The two independent random samples support inference to their respective populations. Each sample is no more than 10% of its population: \(100\le0.10(1000)=100\), and the stated populations are at least 1,000. The sample sizes are therefore no more than 10% of their respective populations. Under the null hypothesis, the pooled sample proportion is \(\hat{p}_{pool}=(60+40)/(100+100)=0.50\). The Large Counts checks use this pooled proportion: in each group, the expected success count is \(100(0.50)=50\), and the expected failure count is also 50. All are at least 10. The conditions support the two-proportion \(z\) test.
Do. The sample proportions are \(\hat{p}_1=60/100=0.60\) and \(\hat{p}_2=40/100=0.40\). The pooled estimate is used in the test statistic because the null hypothesis says the population proportions are equal:
For the two-sided alternative, the p-value is \(2P(Z\ge2.828)\approx0.0047\), rounded to four decimal places. Since \(0.0047<0.05\), we reject \(H_0\).
Conclude. The samples provide convincing evidence that the appointment proportions differ between the two customer populations. The observed proportion is higher in the text-reminder sample than in the email-reminder sample. Because these were random samples rather than a randomized experiment assigning reminder type, the result supports a difference between the populations represented by the samples, not a cause-and-effect claim that text reminders produce more appointments.
Make Each Part Earn Its Place
A complete response is a connected argument, not four unrelated statements. The procedure named in Plan must match the parameter and data structure in State. The condition checks must support that procedure. The Do calculation must use the correct model and match the stated alternative or confidence level. The conclusion must answer the original question rather than merely repeat a number.
Define the population parameter. For a test, give \(H_0\) and \(H_a\); for an interval, identify the parameter being estimated.
Name the inference procedure and verify the relevant conditions using the design and data. Be explicit about any condition that cannot be checked from the information provided.
Show the formula, substitution, and result. For a test, report the test statistic and p-value; for an interval, report the confidence interval and its level.
Answer the question in context. For a test, use “convincing evidence” when appropriate and “fail to reject” when the evidence is insufficient. For an interval, state what population parameter the interval estimates.
A test and an interval give related but different forms of evidence. A p-value measures how surprising results at least as extreme as the observed result would be if the null hypothesis were true. It is not the probability that the null hypothesis is true. A confidence interval gives a range of plausible parameter values under the method’s conditions; it is not a range for individual observations.
Common Mistakes and AP Exam Tips
- Naming only the statistic. Saying “\(\hat{p}=0.60\)” does not identify the population quantity in the question. Define \(p\) or \(\mu\), and name the inference procedure that addresses it.
- Listing conditions without evidence. “Random, independent, and large enough” is too vague. State what was randomly sampled or assigned, show the 10% comparison when it applies, and report the relevant Large Counts or distribution check.
- Using interval conditions for a test, or vice versa. For a proportion test, check Large Counts using the null value (or the pooled proportion for a two-proportion test). For a proportion interval, use the observed sample proportion or proportions, as described in the earlier conditions tutorial.
- Reporting a statistic but not the p-value. A test statistic alone does not tell the reader how much evidence the data provide against the null hypothesis. Report the p-value and compare it with the stated significance level.
- Writing “accept the null.” A large p-value does not prove the null hypothesis. Say “fail to reject \(H_0\)” and describe the lack of convincing evidence for the alternative claim.
- Giving a conclusion without context. “Reject \(H_0\)” is a decision, not a complete contextual conclusion. Name the population or groups and the specific proportion or mean being discussed.
- Overstating what the design allows. Random sampling supports generalizing to the population sampled; random assignment supports cause-and-effect conclusions about assigned treatments. Neither design feature automatically provides both kinds of inference.
- Interpreting an interval as a statement about individuals. A confidence interval estimates a population parameter, such as a mean or proportion. Do not say that a stated percentage of individual values lies inside a confidence interval for the mean.
For full credit, make the reasoning explicit enough that a reader can follow it without guessing. A clear test response states the parameter and hypotheses, names the test, supports the conditions with contextual evidence, shows the statistic and p-value, and concludes in context. A clear interval response does the same for the interval procedure and ends with an interpretation of the population parameter at the stated confidence level.
Check Your Understanding
For each prompt, identify what belongs in a complete response. Pay attention to the parameter, procedure, conditions, calculation, and contextual conclusion.
- A random sample of 80 residents from a town of 2,000 finds that 52 support a proposed park. What population parameter and hypotheses would you state for a two-sided test of whether support differs from 0.60? Which Large Counts values are used for the test?
- A 90% confidence interval for a population mean is reported as \((12.4,15.8)\) hours. Write an interpretation that names the population parameter and avoids describing individual values.
- A two-sided test has p-value 0.12 and uses \(\alpha=0.05\). What decision should the response state, and how should the conclusion be phrased in context?
- Why is “the result proves the treatment caused an increase” not automatically appropriate when two independent random samples, rather than random assignment, were used?
- List the four parts of a complete inference response and state one piece of information that belongs in each part.