Connect Errors, Power, and Recommendations
Questions about errors and power often ask for more than a definition. You may need to identify a possible error, explain how a proposed change affects power, and recommend what a researcher or decision-maker should do. The pieces connect, but they are not interchangeable: an error description is conditional on what is actually true, power is a probability for a specified alternative, and a recommendation weighs statistical goals against practical consequences.
As in “Writing Error Descriptions and Consequences for Free Response,” translate the hypotheses into the situation before naming an error. The previous tutorials on “How Sample Size Affects Power,” “How Significance Level Affects Power and Type II Error,” and “How Effect Size Affects Power” establish how those factors influence power. Here, the new skill is combining those ideas to answer a multi-part question without confusing what the test decides with what someone should do next.
Recall that a Type I error is rejecting a true \(H_0\), while a Type II error is failing to reject a false \(H_0\). For a specified false-null value \(p_1\), power is the probability of rejecting \(H_0\) when \(p_1\) is the true population proportion. A recommendation should respect these conditional meanings. For example, increased power makes it more likely that a test will detect the specified effect; it does not guarantee a statistically significant result or prove that an effect exists.
A Four-Part Approach to Multi-Part Items
A useful approach is to answer each requested part explicitly, then check that your recommendation follows from the scenario rather than from a general slogan such as “more power is always better.”
Identify the population parameter, the null condition, and the direction of the alternative in context.
Describe a Type I or Type II error only under the population truth that makes the test decision wrong.
Say which factor changes, what is held fixed, and how that change affects power for the stated alternative.
Choose a reasonable action, explain its benefit, and acknowledge a relevant cost or trade-off.
The third step matters because power comparisons depend on what is held fixed. At fixed hypotheses and significance level, a larger sample generally increases power. At fixed sample size and significance level, a specified alternative farther from the null value in the direction of \(H_a\) generally has greater power. At fixed sample size and alternative, increasing \(\alpha\) generally increases power but also increases the probability of a Type I error. State the comparison rather than implying that one change has the same effect in every setting.
Worked Example: A Medication Reminder
A clinic tests whether a new text reminder increases the proportion of patients who collect a prescribed medication within seven days. Let \(p\) be the proportion of all eligible patients offered the new reminder who collect the medication within seven days. The hypotheses are \(H_0:p=0.68\) and \(H_a:p>0.68\). The clinic uses \(\alpha=0.05\). Suppose its test fails to reject \(H_0\).
(a) Describe a possible Type II error.
A Type II error occurs if the clinic fails to find convincing evidence that the collection proportion with the reminder is greater than 0.68 when, in fact, the true proportion is greater than 0.68. A possible consequence is that the clinic might stop using a reminder that actually increases timely medication collection.
(b) The clinic is considering a larger sample, with the hypotheses and \(\alpha\) unchanged. Explain the expected effect on power.
For a specified true proportion \(p_1>0.68\), a larger sample generally increases the test’s power. The larger sample makes the test more likely to reject \(H_0\) if that specified increase is real, so it generally lowers the probability of a Type II error at that \(p_1\). It does not change the chosen significance level, which remains the probability of a Type I error for the procedure.
(c) Recommend whether the clinic should consider collecting a larger sample.
If the clinic can recruit more eligible patients without unreasonable cost or delay, collecting a larger sample is a reasonable way to improve the chance of detecting a meaningful increase without raising \(\alpha\). The recommendation is especially relevant if missing a real increase could mean losing an opportunity to improve timely medication collection. The clinic should also consider whether the additional time and resources are justified; a larger sample improves power but cannot guarantee rejection.
Why this answer is complete: It identifies the decision and the population truth needed for a Type II error, states the power comparison with the other factors held fixed, and makes a recommendation tied to both a potential benefit and a practical constraint.
Recommendations Should Reflect the Error Trade-Off
A recommendation is not correct merely because it increases power. The significance level \(\alpha\) is the probability of a Type I error. Raising \(\alpha\) generally increases power for a specified alternative, but it also makes rejecting a true null more likely. If a false alarm would have serious consequences, increasing \(\alpha\) may not be a sensible way to improve power. The earlier tutorial “Choosing Alpha Based on Error Consequences” develops this trade-off.
A larger sample can often increase power without raising \(\alpha\), but it may require more time, money, or participants. A researcher should consider whether a sample large enough to detect a meaningful effect is feasible. The size of the effect also matters: a test tends to have greater power for an alternative farther from the null, but researchers should not describe an effect as likely or important simply because it would be easier to detect.
Worked Example: A Community Recycling Pickup
A town tests whether a new pickup schedule increases the proportion of households that set out recyclable material on collection day. Let \(p\) be the proportion of households in the town that would set out material under the new schedule. The hypotheses are \(H_0:p=0.50\) and \(H_a:p>0.50\). The town has a fixed sample size and uses \(\alpha=0.05\). Consider two possible true proportions for planning: \(p_1=0.54\) and \(p_1=0.62\).
(a) At the same sample size and significance level, for which specified alternative would the test generally have greater power?
The test would generally have greater power when \(p_1=0.62\). That value is farther above the null value of 0.50 than 0.54 is, so the difference between the null model and this specified alternative is larger. The comparison is about power for those particular true proportions, not a claim about which proportion will actually occur.
(b) The town proposes using a larger sample while keeping the same hypotheses and \(\alpha\). Explain the effect.
For either specified alternative, a larger sample generally increases power. With more observations, the test is generally more likely to detect a true increase of that size. Because \(\alpha\) stays at 0.05, this proposal improves power without increasing the test’s significance level.
(c) Recommend a planning choice.
The town should plan around an increase it considers practically meaningful and determine whether it can recruit a sufficiently large sample to have a good chance of detecting that increase. If the town is mainly concerned about a modest increase such as a true proportion of 0.54, it should recognize that this is closer to the null value and generally harder to detect than a proportion of 0.62. It should not claim that a larger effect is more likely simply because power is greater for it. The recommendation should balance the value of detecting a useful increase against the cost of collecting more observations.
Why this answer is complete: It compares effect sizes relative to the null, separately explains the role of sample size, and bases the recommendation on a meaningful target and feasibility rather than treating power as proof.
Work Through the Decision Before the Recommendation
For a multi-part exam item, avoid jumping straight to what the organization should do. First establish what a rejection would mean and what a failure to reject would mean. Then identify the relevant error and power change. This order makes it easier to give a recommendation that does not overstate what the test can establish.
In particular, failing to reject \(H_0\) is not proof that the null value is true. It means the data did not provide convincing evidence for the stated alternative at the chosen significance level. A Type II error is a possible explanation if the null is actually false; it is not something we can identify as having occurred from a failure to reject alone.
Worked Example: A Library’s Online Reservation Tool
A library tests whether offering an online reservation tool increases the proportion of borrowed equipment returned on time. Let \(p\) be the proportion returned on time among borrowers offered the tool. The hypotheses are \(H_0:p=0.80\) and \(H_a:p>0.80\). The library’s initial test fails to reject \(H_0\). It is considering either raising \(\alpha\) from 0.05 to 0.10 or keeping \(\alpha=0.05\) and collecting a larger sample.
(a) Describe a possible Type II error and a possible consequence.
A Type II error occurs if the library fails to find convincing evidence that the proportion returned on time with the tool exceeds 0.80 when, in fact, the true proportion is greater than 0.80. As a result, the library could decide not to continue or expand a tool that actually improves timely returns.
(b) Compare the two proposals in terms of power and Type I error.
At the same sample size and for the same specified alternative, raising \(\alpha\) from 0.05 to 0.10 generally increases power and decreases the probability of a Type II error. It also increases the probability of a Type I error: the library would have a greater chance of concluding that the tool increases the population proportion when it does not.
Keeping \(\alpha=0.05\) and collecting a larger sample generally increases power for a specified alternative without raising the significance level. It may, however, take more time or resources. Neither proposal guarantees that the test will reject \(H_0\).
(c) Recommend one proposal if a false claim of improvement could lead to a costly system-wide purchase.
Keeping \(\alpha=0.05\) and seeking a larger sample is the more defensible proposal if the library has the resources to do so. The larger sample can improve power while preserving the lower Type I error probability associated with \(\alpha=0.05\). Raising \(\alpha\) could make a false claim of improvement more likely, which matters if that claim might prompt an expensive purchase. If a larger sample is not feasible, the library should explain the limitations of the evidence rather than implying that failure to reject proves the tool has no effect.
Four-part reasoning: The test seeks an increase above 0.80; the failure to reject could be a Type II error if an increase is real; a larger sample generally raises power without changing \(\alpha\); and the recommendation reflects the stated cost of a false positive as well as the practical cost of gathering more data.
Common Mistakes and AP Exam Tip
- Calling a possible error a known error: A test result does not reveal the population truth. Say that a Type II error “would occur if” the alternative condition were actually true and the test failed to reject.
- Writing “accept \(H_0\)” after a failure to reject: Say “fail to reject \(H_0\)” or “fail to find convincing evidence for \(H_a\).” Do not claim that the test proved the null condition.
- Saying power always increases when \(\alpha\) increases, without the comparison: Specify that sample size and the alternative are held fixed. Also state the cost: a higher \(\alpha\) increases the chance of a Type I error.
- Describing power without naming the alternative: Power is tied to a specified false-null value. “The test has high power” is incomplete if the question asks about power for a particular population proportion.
- Confusing effect size with evidence from the sample: A larger difference between a specified \(p_1\) and \(p_0\) generally means greater power, with other factors fixed. It does not mean the sample has shown that the larger effect is true.
- Giving a recommendation with no reason: Connect the action to a goal and a trade-off. For example, a larger sample can improve power while keeping \(\alpha\) unchanged, but it may require additional resources.
- Overstating a practical consequence: Say an incorrect decision “could” lead to a consequence. Do not claim the test proves the cause of an outcome or that a possible consequence is certain.
For full-credit communication, answer each part in context. A strong response names the test decision, states the population truth that makes it an error, explains the power comparison with the relevant factors held fixed, and justifies the recommendation using a consequence or constraint from the scenario.
Check Your Understanding
For each scenario, explain the requested error or power comparison and make a recommendation when asked.
- A city tests whether a new bus alert increases the proportion of riders who arrive at a stop on time. The test fails to reject \(H_0\). Describe a possible Type II error in context.
- For a one-proportion test with fixed hypotheses and \(\alpha\), explain how increasing the sample size generally affects power for a specified alternative.
- A test uses \(H_0:p=0.40\) and \(H_a:p>0.40\). With the sample size and \(\alpha\) fixed, compare the general power for \(p_1=0.43\) and \(p_1=0.60\).
- A company can either raise \(\alpha\) or collect a larger sample. If a false positive could prompt a costly product launch, which option might better preserve the Type I error rate while improving power? Explain.
- Why is “the test failed to reject, so the null hypothesis is true” an incorrect statement?