Tutorials › AP Statistics › Choosing Between Binomial and Other Models

Probability model interpretation · Tutorial 384 of 1000

Choosing Between Binomial and Other Models

Use the type of variable and the structure of the chance process to choose a suitable probability model before calculating.

Intermediate 10 min read

What You'll Learn

  • Distinguish a count of successes from other discrete variables and continuous measurements.
  • Recognize when a binomial model is a special kind of discrete probability distribution.
  • Decide when a general discrete distribution is more appropriate than a binomial model.
  • Identify when a normal model is a reasonable description of a quantitative measurement.
  • Explain why the variable’s type and the chance process both matter when choosing a model.
  • Use context and stated assumptions to justify a model choice.

Start with the Variable, Then Check the Process

In Assumptions Behind a Probability Model, we considered why a probability model must match the chance process it describes. Now we add a useful first decision: What kind of variable are we modeling? A variable that counts successes in repeated trials calls for a different model from a variable that records a measurement such as time or temperature.

Three model choices often appear in introductory probability: a binomial distribution, a general discrete distribution, and a normal distribution. A binomial distribution is a particular kind of discrete distribution, but it has a specific structure. A general discrete distribution can describe discrete numerical outcomes without that structure. A normal distribution is a continuous model, often used for quantitative measurements.

Definition: A discrete random variable has a countable set of possible values, such as \(0,1,2,\ldots\). A continuous random variable can take any value within an interval, such as a measurement recorded to greater or lesser precision. A probability model should fit both the variable’s possible values and the process producing them.

The word “countable” includes a finite list of values and an unlimited sequence such as the nonnegative integers. By contrast, a measurement like elapsed time can, in principle, take any value in an interval. A device may display time rounded to the nearest minute, but the underlying time is still a measurement on a continuous scale.

Three Model Choices

A binomial model applies when the random variable counts successes across a fixed number of trials. As in Defining n p and X in Context, define the trial, success, and random variable before selecting the model. Then check the BINS conditions: binary outcomes, independence, a fixed number of trials, and the same probability of success on each trial. Earlier tutorials explain these conditions and the 10% condition when sampling without replacement.

A general discrete distribution is appropriate when the variable has distinct possible numerical values with probabilities assigned to them, but the variable does not have the binomial structure. For example, a model might directly give the probabilities of zero, one, two, or three interruptions in an hour. It does not become binomial merely because its values are whole-number counts.

A normal model is appropriate for a quantitative variable whose distribution is reasonably described by a continuous, symmetric, single-peaked curve. A normal model assigns probabilities to intervals, not to individual exact values. This matches measurements such as duration, length, or temperature more naturally than a count of successes.

Model-selection checklist:
  • Is the variable a count of successes? If so, check whether there is a fixed number of binary trials, independence, and a constant success probability. If all apply, a binomial model may fit.
  • Is it discrete but not a binomial success count? Use a general discrete distribution if the possible values and their probabilities are specified or can be modeled directly.
  • Is it a measurement on a continuous scale? Consider a normal model if the distribution’s shape and context support it.
  • Does the process support the assumptions? A model’s mathematical form is not enough; its assumptions must be reasonable for the situation.

These choices are not simply labels for “small,” “medium,” and “large” numbers. They describe different structures. A binomial variable has a particular trial-based explanation. A general discrete variable needs possible values and probabilities but not that explanation. A normal variable is modeled with a continuous curve and parameters for its center and spread.

Worked Example: A Count That Is Binomial

Worked Example: A Count That Is Binomial

A fictional quality-control process checks 12 independently produced seals. For each seal, the probability it fails inspection is 0.08, and the process is assumed to have the same failure probability for every seal. Let \(X\) be the number of seals that fail. Choose a model and find the probability that none fail.

State. \(X\) is the number of failed seals among 12 checked seals. Its possible values are the whole-number counts from 0 through 12.

Plan. Check the binomial conditions. There is a fixed number of 12 trials. Each seal has two relevant outcomes, fail or pass. The trials are stated to be independent, and the failure probability is the same, \(p=0.08\), on every trial. Therefore a binomial model is appropriate.

Do. In the notation from Writing Binomial Probabilities in Proper Notation, \(X\sim B(12,0.08)\). “None fail” means \(X=0\). Use the exact binomial probability:

$$ P(X=0)=\binom{12}{0}(0.08)^0(0.92)^{12} =1\cdot1\cdot(0.92)^{12} \approx0.3677 $$

The calculator command \(\operatorname{binompdf}(12,0.08,0)\) gives approximately 0.3677, rounded to four decimal places. This is a probability for the count \(X\), not a normal-curve area.

Conclude. Under the stated process and binomial model, the probability that none of the 12 seals fail inspection is about 0.3677.

The values of \(X\) are discrete, but that fact alone does not establish a binomial model. The fixed number of trials and the other conditions make the binomial choice appropriate here. If the failure chances differed among seals or the outcomes were dependent, the same binomial calculation would not be justified without further argument.

Worked Example: A Discrete Count Without Binomial Trials

Worked Example: A Discrete Count Without Binomial Trials

A fictional transit analyst uses a probability model for \(X\), the number of service interruptions on a particular route during one hour. The model assigns these probabilities:

Interruptions, \(x\)0123
Probability, \(P(X=x)\)0.550.300.120.03

Choose a model type and find the probability of at least two interruptions.

State. \(X\) is a discrete numerical count with possible values 0, 1, 2, and 3.

Plan. First verify that the listed values and probabilities can describe a distribution. Then choose between binomial and general discrete models. The scenario provides a probability for each possible count, but it does not describe a fixed number of independent trials with a common probability of “success.” Thus the direct probability table is a general discrete distribution, not a binomial model.

Do. The probabilities are nonnegative and sum to 1:

$$ 0.55+0.30+0.12+0.03=1.00 $$

“At least two” means \(X=2\) or \(X=3\). These outcomes do not overlap, so add their probabilities:

$$ P(X\geq2)=P(X=2)+P(X=3) =0.12+0.03=0.15 $$

The complement provides a check: \(P(X<2)=0.55+0.30=0.85\), and \(1-0.85=0.15\).

Conclude. The listed probabilities form a general discrete distribution, and the model assigns probability 0.15 to at least two service interruptions during the hour.

A count can have discrete values without being generated by a fixed set of binary trials. The table describes the probabilities directly; inventing a number of trials or a per-trial probability would add assumptions not supplied by the situation.

Worked Example: A Measurement Suited to a Normal Model

Worked Example: A Measurement Suited to a Normal Model

In a fictional battery-testing process, the operating time \(X\) for a randomly selected battery is modeled as normal with mean 8.4 hours and standard deviation 0.9 hour. Find the probability that a battery operates between 7.5 and 9.0 hours.

State. \(X\) is the operating time, in hours, of one randomly selected battery. It is a quantitative measurement, and the situation explicitly gives a normal model.

Plan. Use the continuous normal model \(X\sim N(8.4,0.9)\), where the parameters are the mean and standard deviation, as in Notation N(mu, sigma) and Parameters. This is not a count of successes, so a binomial model is not the natural choice. The normal model is stated for the battery times; if it were not stated, evidence about the shape of the distribution would be needed to justify it.

Do. Enter the interval limits, mean, and standard deviation in \(\operatorname{normalcdf}\):

$$ P(7.5\leq X\leq9.0) =\operatorname{normalcdf}(7.5,9.0,8.4,0.9) \approx0.5889 $$

As a check using standardized bounds, \(z=(7.5-8.4)/0.9=-1\), and \(z=(9.0-8.4)/0.9\approx0.67\). The area between those z-scores is about 0.5889, consistent with the calculator result, rounded to four decimal places.

Conclude. Under the stated normal model, the probability that a randomly selected battery operates between 7.5 and 9.0 hours is about 0.5889.

The model describes the underlying measurement as continuous. Even if a display reports operating time rounded to the nearest tenth of an hour, the recorded decimals do not turn the process into a count of successes. The variable’s meaning remains a measured duration.

Do Not Confuse a Variable’s Type with Its Model

A variable’s type narrows the choices but does not always settle them. Whole-number counts are discrete, yet some are binomial and others need a general discrete distribution. A measured quantity may be continuous, but a normal model is not automatically appropriate just because the variable is a measurement. Check the shape or use a normal model only when the scenario supports it.

There is also a special relationship between binomial and normal models. As discussed in Normal Approximation to the Binomial, a normal distribution can approximate a binomial count when the Large Counts condition holds: \(np\geq10\) and \(n(1-p)\geq10\). That is an approximation to a discrete binomial distribution, not a reason to call the original count continuous or to forget its binomial structure. If the question asks for an exact binomial probability, use the binomial model; if it asks for a normal approximation and the condition is met, use that approximation appropriately.

The exact variable also matters. “Number of customers who make a purchase among 20 visitors” can be binomial if the trials and probability conditions are reasonable. “Number of customers arriving in an hour” is a discrete count, but it is not automatically binomial: the prompt would need to describe a fixed number of trials with binary outcomes and the relevant conditions. A direct table of possible arrival counts and probabilities would instead be a general discrete model.

Worked Example: The Same Topic, a Different Model Choice

Worked Example: The Same Topic, a Different Model Choice

A fictional school club records the number of messages members receive during an evening. A model lists \(X=0,1,2,3,4\) with respective probabilities 0.10, 0.25, 0.35, 0.20, and 0.10. A student says, “This must be binomial because messages either arrive or do not arrive.” Evaluate the model choice and find \(P(X\leq1)\).

State. \(X\) is a discrete count of messages received during an evening, with the possible values and probabilities given in the model.

Plan. A message either arriving or not arriving during a tiny time interval does not, by itself, make the stated variable a binomial count. The model does not specify a fixed number of trials, a success probability for each trial, or independent trials. Since probabilities are given directly for the possible values of \(X\), use the general discrete distribution to answer the question.

Do. Check the probability total:

$$ 0.10+0.25+0.35+0.20+0.10=1.00 $$

“At most one” means \(X=0\) or \(X=1\). Therefore:

$$ P(X\leq1)=P(X=0)+P(X=1) =0.10+0.25=0.35 $$

As a complement check, the probability of two or more messages is \(0.35+0.20+0.10=0.65\), and \(1-0.65=0.35\).

Conclude. The supplied table is a general discrete model, and it gives probability 0.35 that a member receives at most one message. The description does not provide the trial structure needed to justify a binomial model.

Common Mistakes and AP Exam Tips

  • Calling every count binomial. A count is discrete, but a binomial model additionally requires fixed binary trials, independence, and the same success probability. Full-credit reasoning names the count and checks those features.
  • Calling every discrete model binomial. A probability table for possible counts can be a general discrete distribution. Do not add an imagined trial structure that the context does not support.
  • Calling every measurement normal. A measurement has the right variable type for a continuous model, but normality still needs to be stated or supported by the distribution’s shape and context.
  • Treating a normal approximation as a new exact model. When approximating a binomial distribution, identify it as an approximation and check both parts of the Large Counts condition.
  • Choosing by number size alone. A count of only a few outcomes may be binomial, and a count with many possible values may still be general discrete. The generating process, not just the range, determines the model.

For a complete AP response, define the random variable in context, identify the relevant type of variable, justify the model using the process or stated probability structure, and then calculate. If information is missing, say which assumption cannot be checked rather than claiming that the model is certain.

AP Exam Tip: Ask two questions in order: “What values can the random variable take?” and “What process or probability information supports this model?” A count is not automatically binomial, and a measurement is not automatically normal.

Key Takeaway

Model selection begins with the nature of the random variable, then uses the chance process to choose more precisely. Binomial models describe counts of successes under specific trial conditions; general discrete models describe other discrete outcomes with assigned probabilities; normal models describe suitable continuous measurements or, under the Large Counts condition, approximate some binomial distributions.

Key takeaway: Match the model to both the values the variable can take and the process that produces them. Use a binomial model only for a qualifying success count, a general discrete model for other specified discrete probabilities, and a normal model when a continuous model is justified.

Check Your Understanding

For each situation, identify the most appropriate model type and explain the key reason for your choice.

  1. A researcher counts how many of 15 randomly selected seedlings sprout. Each has two outcomes, and the problem states that the trials are independent with the same sprouting probability. Which model is appropriate?
  2. A probability table gives the chances of 0, 1, 2, or 3 power interruptions during a day, but describes no fixed set of independent binary trials. Which model fits the table?
  3. A machine measures the time, in seconds, that a package takes to pass along a conveyor. The times are described as roughly symmetric and single-peaked. Which model might be reasonable, and what evidence supports it?
  4. A student says that the number of emails received is binomial because each email either arrives or does not arrive. What additional structure would be needed to justify that claim?
  5. A binomial count satisfies \(np\geq10\) and \(n(1-p)\geq10\), and a problem requests a normal approximation. What should the student check and explain before using the normal model?