Tutorials › AP Statistics › Common Errors with Binomial Distributions

Binomial distributions · Tutorial 359 of 1000

Common Errors with Binomial Distributions

Practice auditing binomial problems before calculating so that the model, event, and calculator inputs match the situation.

Intermediate 11 min read

What You'll Learn

  • Translate “fewer than,” “at most,” and “between” into precise binomial events and inclusive calculator bounds.
  • Identify \(p\) from the outcome the random variable counts, not from a tempting but different outcome.
  • Check the binomial conditions before using a binomial probability command.
  • Use an event-translation audit to catch endpoint and complement mistakes.
  • Explain why a calculator result does not make an unjustified binomial model valid.

Why Binomial Answers Go Wrong

Once a problem has been identified as binomial, calculating a probability may take only a few calculator steps. But a correct calculator command cannot rescue a mismatched model or event. Common errors often begin before the calculation: the success outcome is defined incorrectly, a condition is skipped, or the wording is translated into the wrong count.

In Defining \(n\), \(p\), and \(X\) in Context, you learned to name the trial, define success, and state what \(X\) counts. In Using binomcdf for At Most Probabilities, Finding At Least Probabilities with binomcdf, and Probabilities Between Two Values for Binomial Variables, you learned how cumulative commands handle their endpoints. This tutorial uses those ideas as an error-checking toolkit rather than introducing a new probability rule.

Key technique: Before calculating, do an event-translation audit: write what \(X\) counts, express the requested event with an inequality or equality, identify the relevant \(n\) and \(p\), and check that the binomial conditions support the model. Only then choose a calculator command.

Audit the Event Before Choosing a Command

A binomial cumulative probability includes all possible counts from 0 through its final input. Thus, \(\operatorname{binomcdf}(n,p,k)\) is \(P(X\leq k)\), not \(P(X<k)\). This inclusive endpoint is a frequent source of off-by-one errors. Words that sound similar can require different cutoffs: “at most 4” includes 4, while “fewer than 4” does not.

For a range, check both endpoints. “Between 3 and 6, inclusive” means \(3\leq X\leq6\). Its cumulative calculation subtracts the probability through 2 from the probability through 6. Subtracting the cumulative probability through 3 would remove the probability of 3 as well, which is not wanted.

A useful check is to list the integer counts that should be included. If the requested event is “at least 5,” for example, the included counts are 5 and above, not 4 and above. The complement therefore ends at 4. The “minus one” in this complement calculation is deliberate: it keeps the threshold count out of the complement.

Identify the Success Probability from the Definition of \(X\)

The letter \(p\) is the probability of the outcome labeled success on one trial. Success does not have to mean a desirable result. If \(X\) counts defective items, then a defective item is success for this random variable, and \(p\) is the probability of a defect. Using the probability of a good item instead would model a different count.

This is why the definition of \(X\) should come before entering a command. Ask: “What one outcome adds 1 to the count?” The probability of that outcome is \(p\). The probability of the other outcome is \(1-p\). As discussed in The Binomial Probability Formula, a binomial model needs the success probability for the outcome being counted.

Check the Model, Not Just the Arithmetic

The BINS checklist from The BINS Checklist for Binomial Conditions asks whether outcomes are Binary, trials are Independent, the number of trials is fixed, and the probability of success is the Same for each trial. A binomial calculator command assumes these features; it does not check them for you.

When sampling without replacement from a finite population, also consider the 10% condition from the previous tutorial, The 10% Condition for Independence in Binomial Settings. If the condition is not satisfied, it does not support treating the selections as approximately independent. A sample fraction that is small is not a substitute for checking the actual population and sample sizes. Nor does satisfying the 10% condition fix a nonrandom selection process or a changing success probability.

Audit checklist:
  • Define one trial, the success outcome, and the count \(X\).
  • Identify \(n\) and the per-trial success probability \(p\) from the context.
  • Check Binary outcomes, Independence, fixed Number of trials, and the Same \(p\).
  • If relevant, check the 10% condition for sampling without replacement.
  • Translate the requested event, including its endpoints, before choosing a command.

Worked Example: Inclusive Bounds for a Range

Worked Example: Inclusive Bounds for a Range

A game designer models each of 8 independent plays as a success with probability 0.50. Let \(X\) be the number of successful plays. Find the probability of between 3 and 5 successes, inclusive.

State. The model is \(X\sim B(8,0.50)\). The requested event includes 3, 4, and 5 successes, so it is \(3\leq X\leq5\).

Plan. Each play has two outcomes, the plays are independent, there are 8 fixed plays, and the success probability is 0.50 on each play. These satisfy BINS as stated in the model. For the probability, subtract the cumulative probability through 2 from the cumulative probability through 5.

Do. The calculator inputs use inclusive upper cutoffs:

$$ P(3\leq X\leq5) =\operatorname{binomcdf}(8,0.50,5) -\operatorname{binomcdf}(8,0.50,2) $$

The cumulative probabilities are \(219/256=0.85546875\) and \(37/256=0.14453125\), respectively. Therefore:

$$ P(3\leq X\leq5) =0.85546875-0.14453125 =0.7109375 \approx0.7109 $$

As a check, the probabilities of exactly 3, 4, and 5 successes add to \((56+70+56)/256=182/256=0.7109375\). If the lower cumulative cutoff had been 3 instead of 2, the probability of exactly 3 successes would have been incorrectly removed.

Conclude. Under this model, the probability of 3, 4, or 5 successful plays is about 0.7109.

Worked Example: “Fewer Than” Is Not “At Most”

Worked Example: “Fewer Than” Is Not “At Most”

A technician tests 9 independent sensors. Each sensor has probability 0.30 of triggering a warning. Let \(X\) count the sensors that trigger a warning. Find the probability that fewer than 3 trigger a warning.

State. Let \(X\sim B(9,0.30)\), where success means that a sensor triggers a warning. “Fewer than 3” excludes 3, so the event is \(X<3\), equivalently \(X\leq2\).

Plan. The event has a whole-number upper cutoff of 2. Use \(\operatorname{binomcdf}(9,0.30,2)\). Do not use a cutoff of 3, which would include the outcome \(X=3\).

Do. The probability can be calculated with the cumulative command or by adding the probabilities of 0, 1, and 2 warnings:

$$ \begin{aligned} P(X<3)&=P(X\leq2)\\ &=\operatorname{binomcdf}(9,0.30,2)\\ &=0.7^9+9(0.30)(0.7)^8+\binom{9}{2}(0.30)^2(0.7)^7\\ &=0.040353607+0.155649627+0.266827932\\ &=0.462831166\approx0.4628 \end{aligned} $$

The first term represents zero warnings, the second represents one, and the third represents two. The endpoint check confirms that no probability for exactly 3 warnings is included.

Conclude. The probability that fewer than 3 of the 9 sensors trigger a warning is about 0.4628.

Worked Example: Success Means the Outcome Being Counted

Worked Example: Success Means the Outcome Being Counted

In a production process, each of 10 independently produced parts has probability 0.08 of being defective. Let \(X\) be the number of defective parts among the 10. Find the probability of at least one defective part.

State. Define success as “the part is defective,” because \(X\) counts defective parts. Thus \(n=10\), \(p=0.08\), and \(X\sim B(10,0.08)\).

Plan. A part is either defective or not defective; the 10 production outcomes are stated to be independent; the number of parts is fixed at 10; and the defect probability is the same for each part. BINS is satisfied by the scenario. “At least one” means \(X\geq1\), so use the complement of \(X=0\).

Do. The probability of zero defective parts is the cumulative probability through 0. Use the defect probability—not the probability \(0.92\) that a part is not defective—as \(p\) in the model for \(X\):

$$ \begin{aligned} P(X\geq1) &=1-P(X=0)\\ &=1-\operatorname{binomcdf}(10,0.08,0)\\ &=1-(0.92)^{10}\\ &=1-0.4343884542\\ &=0.5656115458\approx0.5656 \end{aligned} $$

Using \(p=0.92\) would instead describe the count of nondefective parts. It would not be the model for the defined random variable \(X\), which counts defects.

Conclude. The probability that at least one of the 10 parts is defective is about 0.5656.

Worked Example: A Calculator Cannot Fix a Failed Condition

Worked Example: A Calculator Cannot Fix a Failed Condition

A school has 300 students, including 45 who participate in a particular club. A staff member randomly selects 40 students without replacement and wants the probability that at least 8 selected students participate. Is a binomial model justified by the stated information and the 10% condition?

State. Let \(X\) be the number of selected students who participate. The proposed success is club participation, so \(p=45/300=0.15\), and \(n=40\).

Plan. Check BINS and the 10% condition before calculating. The outcome is binary and the sample size is fixed. Selection is random, but it is without replacement, so examine whether the condition supports treating selections as approximately independent. The same-probability condition must also be considered in light of that selection process.

Do. The population size is \(N=300\), while ten times the sample size is:

$$ 10n=10(40)=400 \qquad\text{and}\qquad N=300<400 $$

Equivalently, \(n/N=40/300\approx0.1333\), or about 13.33% of the population. The 10% condition is not satisfied. As explained in The 10% Condition for Independence in Binomial Settings, this condition does not support treating these without-replacement selections as approximately independent. In fact, the pool of students and its club-participation proportion change as each student is selected. Although the outcome is binary, \(n\) is fixed, and the sample is random, the binomial assumptions are not adequately supported by this check.

A calculator will still return a number if asked for \(\operatorname{binomcdf}(40,0.15,7)\), the complement that would correspond to “at least 8” under a binomial model. But that output is not justified as the requested probability by the information here. It is a calculation under an unsupported model, not evidence that the model fits the selection process.

Conclude. The 10% condition fails because 300 is less than 400. A binomial model is not justified as an approximate independent-trials model by this condition, so we should not report a binomial calculator result as the answer for this sample.

Common Mistakes and AP Exam Tips

  • Using the cutoff as if it were exclusive. The final input to \(\operatorname{binomcdf}\) is inclusive. For “fewer than 6,” use a cutoff of 5; for “at most 6,” use 6.
  • Subtracting the wrong lower cumulative probability for a range. For \(a\leq X\leq b\), subtract the cumulative probability through \(a-1\), not through \(a\).
  • Using the probability of the opposite outcome as \(p\). First define what \(X\) counts. If \(X\) counts failures, use the failure probability as \(p\), even if success sounds like the more natural label in everyday speech.
  • Checking only the arithmetic. A correct-looking calculator result is not a valid binomial probability if the setting does not support BINS or, when relevant, the 10% condition.
  • Calling a failed condition “close enough” without justification. State which condition fails and avoid presenting the binomial result as justified. A calculator cannot decide whether an assumption is reasonable.
  • Reporting a number without its event or context. Connect the probability to the defined count and the requested outcome so a reader can tell what the number means.

For full-credit communication, define \(X\), identify \(n\) and \(p\), and write the event before showing the calculator command. For a cumulative calculation, make the inclusive cutoff visible. For a model check, state which conditions hold and which do not; when sampling without replacement, explain what the 10% check says about using a binomial approximation.

AP Exam Tip: Do not write only “I used binomcdf.” Write the event, such as \(P(X\leq5)\), show the matching command, and interpret the result in context. If a condition fails, say so before using or rejecting a binomial calculation.

Key Takeaway

Most binomial mistakes can be caught by separating three decisions: whether the model is appropriate, what outcome \(X\) counts, and which counts satisfy the question. Make those decisions before entering values in a calculator.

Key takeaway: Check the binomial conditions, define success to match \(X\), and translate the wording into exact inclusive bounds. Then choose a command whose inputs match that event and interpret its result in context.

Check Your Understanding

For each item, identify the event or modeling issue before calculating.

  1. If \(X\) is binomial and the question asks for “at most 4,” what inequality describes the event, and what is the final input to \(\operatorname{binomcdf}\)?
  2. A question asks for “fewer than 4.” Which cumulative cutoff should be used, and why is the cutoff not 4?
  3. A binomial variable \(X\) counts orders containing an error. The probability an order contains an error is 0.06. What value of \(p\) belongs in the model for \(X\)?
  4. For \(2\leq X\leq5\), which two cumulative probabilities should be subtracted? Explain the lower cutoff.
  5. A random sample of 50 people is selected without replacement from a population of 400. Does the 10% condition support treating the selections as approximately independent? Show the comparison and state what the check does—and does not—establish.