From Naming Conditions to Justifying a Model
In Common Errors in Evaluating Probability Models, you saw why a correct calculation does not, by itself, establish that a model fits the chance process. This tutorial focuses on the next step: writing the justification clearly. On an AP Statistics question, listing condition names is usually not enough. Connect each relevant condition to details in the setting, then state what those details imply about using the model.
A good justification makes a short chain of reasoning visible: condition, evidence, implication. For example, rather than writing only “independence holds,” describe how the trials are produced and why one result is unlikely to change the chance of another. Then say whether independence seems reasonable for the model.
This frame is not a script to repeat mechanically. The details and conclusion must fit the model and the situation. A condition might be supported by the description, questionable because of a shared influence, or impossible to assess from the information given. Say which is true; do not claim more than the context supports.
What Makes a Justification Convincing?
A useful response does three things. First, it identifies the condition that matters. Second, it points to concrete information about how the outcomes are generated. Third, it explains the consequence for the model. That last link matters: readers should understand why the detail supports or challenges the condition.
| Part of the justification | What to include | Example of useful wording |
|---|---|---|
| Condition | The model’s requirement | “The binomial model requires independent trials.” |
| Evidence | A relevant detail about the process | “The items are selected without replacement from a batch of 5,000.” |
| Implication | How the evidence affects the model choice | “Because the sample is a small fraction of the batch, treating the selections as approximately independent is reasonable.” |
A condition is not supported just because the problem uses a familiar word. “Random” needs context: who or what was selected, and how? “Same probability” needs a reason to think the process stays stable across trials. “Approximately normal” needs evidence about the distribution’s shape. If the question gives no evidence for an assumption, identify that gap rather than inventing support.
For binomial models, use the conditions explained in Common Errors with Binomial Distributions: a fixed number of trials, two outcomes per trial, independent trials, and a constant success probability. For normal models, build on Checking Whether Data Are Approximately Normal and Judging Whether a Normal Model Fits: consider the distribution’s shape and the evidence provided. The goal here is not to re-derive those conditions; it is to communicate how the setting supports them.
Worked Example: Sampling Batteries from a Large Run
Worked Example: Sampling Batteries from a Large Run
A fictional manufacturer models the number of defective batteries in a random sample of 20 batteries from a stable production run of 5,000. Its proposed model is \(X\sim\operatorname{Binom}(20,0.08)\). The 0.08 defect probability comes from the manufacturer’s stated model for this run. Write a justification for using the binomial model.
State. Let \(X\) be the number of defective batteries in the sample of 20. The proposed model is binomial, with \(n=20\) and \(p=0.08\).
Plan. Check the four binomial conditions and connect each one to the stated process. In particular, because batteries are selected without replacement, consider whether the selections can reasonably be treated as independent.
Do. The number of trials is fixed at 20, and each selected battery has one of two outcomes for this question: defective or not defective. The problem describes a random sample, which supports treating the selection process as random. The sample is a small fraction of the production run:
Since 0.4% is less than 10%, the 10% condition supports treating selections without replacement as approximately independent. The proposed model assigns the same defect probability, 0.08, to each battery. The description of a stable production run gives a reason that this probability may be similar across the sampled batteries. It does not prove that every battery has exactly the same chance, but it supports the assumption for this model.
Conclude. A binomial model is reasonable for describing the number of defective batteries in this sample: there are 20 fixed trials, two outcomes per battery, a stated common defect probability, and random sampling of only 0.4% of the run, supporting approximate independence under the 10% condition. The conclusion is about this sample from the specified stable run; it does not establish that the same model applies to other runs or production conditions.
Notice that the conclusion does not say “all conditions are proven.” Random selection and a stable run provide support, while the common probability remains a modeling assumption. AP responses are stronger when they distinguish evidence from certainty.
Worked Example: When Trial Conditions May Change
Worked Example: A Binomial Model for Free Throws
A coach proposes a binomial model for the number of made free throws in 12 consecutive attempts by one player during a late practice session. The player shoots from the same line, but the coach says the player becomes tired and changes shooting technique over the session. Evaluate the model’s appropriateness.
State. Let \(X\) be the number of made shots in 12 attempts. The proposed model is binomial.
Plan. Check the fixed number of trials, two outcomes, independence, and constant success probability. A condition can be questionable even when the other conditions appear to fit.
Do. The number of attempts is fixed at 12, and each attempt has two outcomes for this question: made or missed. But the context gives a reason to doubt a constant success probability. Fatigue and a changing technique could alter the player’s chance of making a shot over the session. Independence is also uncertain: a player’s fatigue or adjustment after one attempt could affect later attempts. Shooting from the same line does not, by itself, establish either constant probability or independence.
Conclude. A binomial model is not well supported for these 12 attempts as described. Although the fixed-trial and two-outcome conditions hold, fatigue and technique changes make a constant success probability questionable, and adjustments during the sequence could connect outcomes. A binomial calculation would describe the result only under assumptions that this context gives reason to doubt.
This example shows why “same setting” and “same probability” are not interchangeable. The player shoots from the same line, but relevant conditions change. A full-credit justification names the changing feature and explains which model requirement it threatens.
Worked Example: Assessing a Normal Model for Package Weights
Worked Example: Assessing a Normal Model for Package Weights
A fictional packaging line produces sealed snack packages. A random sample of 120 packages from one stable shift has a roughly single-peaked, symmetric histogram with no apparent outliers. The sample mean weight is 500 grams and the sample standard deviation is 4 grams. Of the 120 weights, 81 fall within 1 standard deviation of the mean, 113 fall within 2 standard deviations, and 119 fall within 3 standard deviations. Evaluate a normal model for package weights from this shift.
State. Let \(X\) be the weight, in grams, of a randomly selected package from this shift. A normal model is being considered.
Plan. Use the distribution’s shape and the empirical-rule percentages as evidence about whether a normal model is reasonable. Also consider how the sample was selected and the scope of the conclusion.
Do. The random sample supports using the observed packages to assess the shift’s distribution. The histogram is described as roughly single-peaked and symmetric, with no apparent outliers, which is consistent with a normal shape. The observed percentages within 1, 2, and 3 standard deviations are:
These are about 67.5%, 94.2%, and 99.2%, respectively, close to the 68%, 95%, and 99.7% empirical-rule benchmarks. The percentages agree with the shape evidence, so they support, but do not prove, a normal model. The packages also come from one stable shift, which limits how far the conclusion should be extended.
Conclude. A normal model appears reasonable for package weights during this shift because the random sample’s distribution is roughly symmetric and single-peaked, has no apparent outliers, and has percentages within 1, 2, and 3 standard deviations close to the empirical-rule benchmarks. This evidence does not establish a normal model for other shifts or for a period when the packaging process changes.
Writing a Clear Conclusion
The conclusion should answer the question directly, but it should also keep the model’s scope visible. “The model is appropriate” can sound more certain and broader than the evidence allows. Prefer a conclusion that states whether the model appears reasonable for the specified process and names an important qualification.
Say what model is proposed and what the random variable measures, including units when relevant.
For a binomial model, check all four conditions. For a normal model, address the evidence about shape and any other details the question provides.
Use details such as the selection method, sample size relative to the population, a stable process, a shared influence, or the described distribution shape.
Conclude that the model appears reasonable, is questionable, or is not appropriate for the stated situation, and mention limits when they matter.
If a condition cannot be checked from the information given, state that plainly. For example: “The problem does not describe how the sample was selected, so the information provided is not enough to assess whether the selection was random.” That is more defensible than claiming random selection simply because a sample is mentioned.
Common Mistakes and AP Exam Communication
- Listing labels without evidence. “BINS holds” does not explain why. A stronger response gives the fixed number, identifies the two outcomes, uses the sampling design to address independence, and explains why a common probability is plausible or uncertain.
- Repeating “random” without describing the design. Say what was randomly selected or assigned. If the method is not supplied, say that randomness cannot be assessed from the description.
- Claiming independence from a large sample size alone. Explain the selection process. When sampling without replacement, compare the sample size with the population and apply the 10% condition when appropriate.
- Assuming identical conditions guarantee a constant probability. A shared location or equipment does not rule out changes in time, fatigue, or other influences. Describe why the chance should stay similar—or what could make it vary.
- Calling a distribution normal because the mean and standard deviation are known. Parameters describe center and spread, not shape. Use the histogram, normal probability plot, or other distribution evidence provided in the question.
- Using absolute language when the evidence is limited. Context details support judgments; they rarely prove assumptions. “Appears reasonable for this process” is often more accurate than “is definitely valid.”
- Extending the conclusion beyond the stated setting. A model supported for one production run or shift is not automatically appropriate for all future runs. Name the population or process the evidence actually represents.
Key Takeaway
A model justification is an argument, not a checklist. The condition identifies what the model needs; the context detail supports or challenges that requirement; and the conclusion explains whether using the model is reasonable for the stated question.
Check Your Understanding
For each situation, identify a relevant condition, the context evidence you would use, and a suitably qualified conclusion.
- A researcher models the number of people who choose a new drink among 50 visitors randomly selected without replacement from 2,000 visitors. What does the 10% condition say about treating selections as approximately independent?
- A binomial model is proposed for 15 machine-produced parts, but the parts come from three machines with different recent maintenance histories. Which condition may be difficult to justify, and why?
- A normal model is proposed for delivery times. The context gives a mean and standard deviation but no plot or description of the distribution. What can you conclude about the model’s shape condition?
- Rewrite this justification so it connects condition, evidence, and implication: “The trials are independent because they are separate.”
- A model seems reasonable for measurements collected during one stable shift. What should a careful conclusion say about applying it to a different shift?