Read a Conclusion Like a Grader
A conclusion can sound persuasive and still leave out an important part of the statistical reasoning. A student might report the correct decision but never connect it to the research question. Another might name the right claim but say the test “proves” it. Evaluating these responses is easier when you look for specific elements instead of deciding whether the answer simply “sounds right.”
In “Common Mistakes in Writing Conclusions,” you studied errors that weaken a mean-test conclusion. Here, you will use those ideas in a new way: as a grader, identify which parts of a response are correct, award credit for those parts, and describe a precise revision. The practice rubric below is an instructional tool for learning to evaluate conclusions. It is not presented as an official College Board scoring rubric.
Use the test’s stated alternative hypothesis, p-value, and significance level as your reference. As in “Making a Decision From P-Value and Alpha,” reject \(H_0\) when \(p\le\alpha\), and fail to reject \(H_0\) when \(p>\alpha\). Then check whether the student’s wording accurately explains what that decision says about the population mean.
A Six-Element Practice Rubric
For practice, score a conclusion out of six points, awarding one point for each element below. A response can earn credit for a correct element even if another element is missing or wrong. This point-by-point approach helps you explain a score and avoid treating one mistake as if it erased everything the student did correctly.
| Element | What to look for |
|---|---|
| 1. P-value comparison | The response compares the given p-value with the stated \(\alpha\). |
| 2. Decision | The response correctly says to reject or fail to reject \(H_0\). |
| 3. Evidence wording | The response says the data do or do not provide convincing evidence, as appropriate. |
| 4. Parameter and context | The response identifies the population mean or means and the measured variable in the situation. |
| 5. Direction of the claim | The response’s contextual claim matches the direction of \(H_a\), or describes a difference when the alternative is two-sided. |
| 6. Appropriate limits | The response avoids claiming proof, certainty, a result for every individual, or a broader conclusion unsupported by the study design. |
These elements are related, but they are not interchangeable. “Reject \(H_0\)” earns the decision element if it is correct, but it does not automatically earn the evidence wording or contextual claim elements. Similarly, a response may correctly name the population mean and direction but fail to explain whether the evidence is convincing.
When grading, use the hypotheses and study description rather than guessing what the student meant. A phrase like “the average is lower” may be too vague to show whether the student means the sample mean or the population mean. If the intended meaning is unclear, do not silently award credit for wording the student did not provide.
A Repeatable Scoring Method
Identify the parameter, alternative hypothesis, p-value, \(\alpha\), and target population from the problem. These determine what a correct conclusion must say.
For every element, record “earned” or “not earned,” and point to the words that support your judgment. Do not award a point based only on what the student may have intended.
State which element is missing or incorrect. For example, distinguish “the decision is wrong” from “the decision is right, but the response does not name the population mean.”
Write a corrected conclusion that keeps the student’s accurate work, fixes the specific errors, and answers the original question in context.
A grader should also avoid double-penalizing the same wording without reason. If a student says “accept \(H_0\),” that is an incorrect decision for this rubric. The same phrase may also signal a mistaken claim that the null is true, which affects the limits element. But explain both issues specifically; do not simply subtract points twice because the phrase feels generally weak.
Worked Example: A Correct Decision With a Vague Claim
Worked Example: A Correct Decision With a Vague Claim
A fictional greenhouse randomly selects tomato seedlings from a production batch and measures the number of days until each seedling produces its first flower. A one-sample t test evaluates whether the population mean time is less than 42 days. The test result is \(p=0.018\), with \(\alpha=0.05\). A student writes:
“Since \(0.018<0.05\), reject \(H_0\). The plants flower sooner.”
Use the six-element practice rubric. The comparison is correct, so element 1 is earned. The decision is also correct, so element 2 is earned. The student does not say that the data provide convincing evidence, so element 3 is not earned. “The plants flower sooner” does not identify the population mean time or clearly state the comparison with 42 days, so element 4 is not earned. The phrase suggests the less-than direction, but does not clearly state the test’s population claim; award element 5 only if the scoring standard accepts that wording as an unambiguous directional claim. Under this rubric’s requirement for a clear contextual claim, it is not earned. The response does not claim proof or overgeneralize, so element 6 is earned.
That gives a score of 3 out of 6: elements 1, 2, and 6. The score identifies a useful strength—the evidence comparison and decision are right—while showing that the response needs a clearer account of what the test concerns.
A full-credit revision is: “Since \(p=0.018<0.05\), we reject \(H_0\). The data provide convincing evidence that the true mean time until first flowering for tomato seedlings in this production batch is less than 42 days.” This names the population mean and variable, matches the direction of the alternative, and avoids suggesting that every seedling flowers sooner.
Worked Example: “Accept the Null” Is Not a Safe Conclusion
Worked Example: “Accept the Null” Is Not a Safe Conclusion
A fictional library system compares checkout times at two branches. A two-sample t test addresses whether the true mean checkout time at Branch A is greater than the true mean checkout time at Branch B. The supplied test result is \(p=0.087\), with \(\alpha=0.05\). A student writes:
“Because \(0.087>0.05\), accept \(H_0\). There is no difference in the average checkout time between the two libraries.”
Element 1 is earned: the student makes the correct comparison. Element 2 is not earned: when \(p>\alpha\), the correct decision is to fail to reject \(H_0\), not to accept it. Element 3 is not earned, because the student does not accurately describe the lack of convincing evidence for the alternative. “There is no difference” treats a failure to reject as proof of equality. Element 4 is earned: the response refers to average checkout time at the two branches, which identifies the relevant quantity and context. Element 5 is not earned: the alternative asks whether Branch A’s mean is greater, while the student asserts no difference. Element 6 is not earned because the response makes an unsupported claim that there is no difference.
The score is 2 out of 6: elements 1 and 4. A key grading distinction is that correctly writing “\(0.087>0.05\)” does not make the following conclusion correct. The comparison and decision are separate rubric elements.
A corrected conclusion is: “Since \(p=0.087>0.05\), we fail to reject \(H_0\). The data do not provide convincing evidence that the true mean checkout time at Branch A is greater than the true mean checkout time at Branch B.” This does not claim that the means are equal; it reports what the test does and does not provide evidence for.
Worked Example: Correct Numbers, Overstated Result
Worked Example: Correct Numbers, Overstated Result
A fictional technology team randomly selects 40 employees from a company of 400 employees. In a randomized experiment, each selected employee is assigned to complete a routine task using either a new scheduling app or the usual method. The researchers compare the population mean completion time for the two methods. The test’s alternative is that the mean time with the app is lower. The supplied result is \(p=0.004\), with \(\alpha=0.01\). A student writes:
“Since \(0.004<0.01\), reject \(H_0\). The test proves that the app lowers the mean completion time for the 40 employees.”
Element 1 is earned because the p-value is correctly compared with \(\alpha\). Element 2 is earned because the correct decision is to reject \(H_0\). Element 3 is not earned: “proves” is not appropriate evidence wording. Element 4 is not earned under this rubric, because the conclusion refers to the 40 employees rather than clearly describing the population of company employees represented by the random sample. Element 5 is earned: the claim that the app lowers the mean completion time matches the direction of the alternative. Element 6 is not earned: the response claims proof and does not carefully state the scope of the conclusion.
The score is 3 out of 6: elements 1, 2, and 5. The student’s p-value comparison and decision are correct, and the direction is right. Those strengths do not cancel the overstatement or the unclear population statement.
A better conclusion is: “Since \(p=0.004<0.01\), we reject \(H_0\). The data provide convincing evidence that, for employees at this company, using the scheduling app results in a lower true mean task-completion time than using the usual method.” Random assignment supports a causal interpretation of the treatment comparison; the random sample supports generalizing to the company’s employees. The conclusion concerns a difference in population means, not a guarantee that every employee will finish faster.
Common Grading Mistakes and AP Exam Tips
- Scoring the answer’s apparent intent: Give credit for what the student actually states. If the response says only “the average is lower,” check whether it identifies the population mean and the comparison; do not fill in missing context for the student.
- Combining the comparison and decision: A correct p-value comparison does not guarantee a correct decision. Score each separately, especially when a student writes “accept \(H_0\)” after correctly noting that \(p>\alpha\).
- Treating “no difference” as equivalent to failing to reject: A large p-value does not establish that population means are equal. A full-credit response says the data do not provide convincing evidence for the stated alternative.
- Accepting “significant” as a complete conclusion: “The result is significant” does not identify the population parameter or answer the research question. Look for a contextual statement about the mean or difference in means.
- Ignoring the alternative’s direction: Compare the student’s claim with \(H_a\). A conclusion about a lower mean does not match a test whose alternative claims a higher mean, even if the p-value comparison is correct.
- Overlooking unsupported scope or certainty: “Proves,” “everyone,” and claims about populations not represented by the study can make a conclusion misleading. Check the sampling and assignment design before crediting generalization or causation.
A useful grader’s note names the element, the evidence in the response, and the needed change. For example: “The decision is correct because the p-value exceeds \(\alpha\), but replace ‘accept \(H_0\)’ with ‘fail to reject \(H_0\),’ and state that the data do not provide convincing evidence for the directional claim.” This feedback is more useful than “be more precise” because it tells the student what to keep and what to revise.
Check Your Understanding
Use the six-element practice rubric to evaluate each response. Explain which elements are earned and suggest a correction where needed.
- A test of whether a population mean exceeds 18 has \(p=0.03\) and \(\alpha=0.05\). A student writes, “Reject \(H_0\). The mean is above 18.” What contextual information or evidence wording could improve the response?
- A two-sided test comparing two population means has \(p=0.22\) and \(\alpha=0.05\). Why does “the population means are equal” go beyond the test result?
- A student correctly writes \(p=0.006<0.01\) and “reject \(H_0\),” but then says the result “proves that every product lasts longer.” Identify two possible rubric problems.
- Why should a grader score the p-value comparison and the reject-or-fail-to-reject decision separately?
- What information about a study’s design helps determine whether a conclusion may generalize to a population or describe a causal effect?