Tutorials › AP Statistics › Matching a Model to a Situation Using Variable Type

Probability model interpretation · Tutorial 385 of 1000

Matching a Model to a Situation Using Variable Type

Match counts, measurements, and proportions to probability models by checking both the variable’s type and the process that produces it.

Intermediate 9 min read

What You'll Learn

  • Distinguish a count, a measurement, and a sample proportion by the values each can take.
  • Identify when a count is suited to a binomial model rather than a general discrete probability table.
  • Explain why a measurement can fit a normal model when its distribution is reasonably normal.
  • Connect a sample proportion to its underlying binomial count and possible values.
  • Justify a normal approximation for a binomial count or sample proportion using the Large Counts condition.

Match the Variable and the Process

In Choosing Between Binomial and Other Models, we compared binomial, general discrete, and normal models. Here we make that choice more systematic by focusing on the variable’s type: Is it a count, a measurement, or a proportion? Then we ask what process produces its values. Neither the variable’s name nor the fact that its values are numbers is enough by itself to identify a suitable model.

The variable’s possible values give an important first clue. A count takes whole-number values. A measurement can, in principle, take any value across an interval. A sample proportion is a ratio formed from a count and a fixed sample size, so its possible values are restricted even though it is written as a decimal. The process and any supplied probability information help determine which model describes those values.

Definition: A general discrete probability distribution lists countable possible values of a random variable and assigns a probability to each. A binomial distribution is a particular discrete model for a count of successes in a fixed number of trials meeting the binomial conditions. A normal distribution is a continuous model used for suitable measurements and, under the Large Counts condition, as an approximation to certain discrete distributions.

A useful distinction is between the original variable and a model used to approximate its distribution. For example, a count remains discrete even when a normal model is used to approximate its probabilities. The approximation does not change what the variable counts.

A Practical Matching Process

Start by defining the random variable in context. Then list or describe its possible values. A number of successes among a fixed set of trials can be binomial if the trial conditions hold. A different count can be represented by a general discrete table if its possible values and probabilities are supplied. A quantitative measurement may fit a normal model when the distribution is reasonably symmetric and single-peaked, or when a normal model is explicitly given.

For a proportion, identify the count and the denominator. If \(\hat{p}\) is the proportion of successes in a fixed sample of size \(n\), then \(\hat{p}=X/n\), where \(X\) is the number of successes. Thus \(\hat{p}\) can take only the values \(0,1/n,2/n,\ldots,1\). If \(X\) is binomial, its probabilities determine the probabilities for \(\hat{p}\). With a sufficiently large sample under the Large Counts condition, a normal model can approximate the distribution of \(\hat{p}\), as discussed in Normal Models for Sample Means and Proportions.

1
Define the variable.
Say exactly what \(X\), a measurement, or \(\hat{p}\) represents and include the units or denominator.
2
Identify its possible values.
Decide whether they are whole-number counts, values on a measurement scale, or proportions restricted by a fixed sample size.
3
Check the model information.
For a binomial count, check the trial conditions. For a discrete table, use the listed values and probabilities. For a normal model, look for a stated model or evidence that its shape is appropriate.
4
State any approximation.
If using a normal model to approximate a binomial count or sample proportion, verify the Large Counts condition and identify the result as an approximation.

This process avoids a common shortcut: choosing a model just because the values “look like” the values of a familiar distribution. Model choice should reflect both what the variable can be and how the chance process generates it.

Worked Example: A Fixed-Trial Count

Worked Example: A Fixed-Trial Count

A fictional greenhouse tests 8 seeds. Each seed either sprouts or does not sprout. The setup states that sprouting outcomes are independent and that each seed has probability 0.25 of sprouting. Let \(X\) be the number of seeds that sprout. Choose a model and find the probability that exactly 2 sprout.

State. \(X\) counts sprouting seeds among the 8 tested. Its possible values are the whole numbers 0 through 8.

Plan. A count is discrete, but that alone does not establish a binomial model. Here there is a fixed number of trials, each trial has two outcomes, the outcomes are stated to be independent, and the sprouting probability is the same, \(p=0.25\), for each seed. These satisfy the binomial conditions, so use a binomial model.

Do. As in Writing Binomial Probabilities in Proper Notation, \(X\sim B(8,0.25)\). The event “exactly 2 sprout” is \(X=2\). Using the binomial formula from The Binomial Probability Formula:

$$ P(X=2)=\binom{8}{2}(0.25)^2(0.75)^6 =28(0.0625)(0.1779785) \approx0.3115 $$

The calculator command \(\operatorname{binompdf}(8,0.25,2)\) gives approximately 0.3115, rounded to four decimal places. The variable is a count, and the trial information—not just its whole-number values—supports the binomial model.

Conclude. Under the stated model, the probability that exactly 2 of the 8 seeds sprout is about 0.3115.

Worked Example: A Discrete Table for a Count

Worked Example: A Discrete Table for a Count

A fictional community center uses a probability model for \(X\), the number of equipment-repair requests it receives during one evening. The model gives:

Requests, \(x\)0123
Probability, \(P(X=x)\)0.200.350.300.15

Choose a model type and find the probability of at least two requests.

State. \(X\) is a discrete count with possible values 0, 1, 2, and 3.

Plan. The table assigns probabilities directly to each possible count. The scenario does not describe a fixed number of independent binary trials with a common probability of success, so the information does not support a binomial model. Use the general discrete distribution represented by the table.

Do. First check that the listed probabilities are nonnegative and sum to 1:

$$ 0.20+0.35+0.30+0.15=1.00 $$

“At least two” means \(X=2\) or \(X=3\). These possible values are mutually exclusive, so add their probabilities:

$$ P(X\geq2)=P(X=2)+P(X=3) =0.30+0.15=0.45 $$

As a check, \(P(X<2)=0.20+0.35=0.55\), so the complement is \(1-0.55=0.45\).

Conclude. The table is a valid general discrete probability model, and it assigns probability 0.45 to receiving at least two repair requests that evening.

The table’s count values do not imply that each request is a “success” in a binomial experiment. A binomial explanation would require a fixed number of trials and the relevant trial conditions. When a problem provides a discrete table instead, use the probabilities it gives rather than inventing a trial structure.

Worked Example: A Measurement Suited to a Normal Model

Worked Example: A Measurement Suited to a Normal Model

In a fictional packaging process, the time \(X\) for a package to pass through a scanner is modeled as normal with mean 42 seconds and standard deviation 5 seconds. Find the probability that a randomly selected package takes between 37 and 49 seconds.

State. \(X\) is the scanner time, in seconds, for one randomly selected package. It is a measurement that can take values across an interval, rather than a count of successes.

Plan. The situation explicitly supplies a normal model, so use \(X\sim N(42,5)\), with mean 42 seconds and standard deviation 5 seconds. If the model were not given, the measurement type alone would not establish normality; evidence that the distribution’s shape is reasonably symmetric and single-peaked would be relevant.

Do. Use the interval limits, mean, and standard deviation in \(\operatorname{normalcdf}\):

$$ P(37\leq X\leq49) =\operatorname{normalcdf}(37,49,42,5) \approx0.7606 $$

As a check, the standardized endpoints are \(z=(37-42)/5=-1\) and \(z=(49-42)/5=1.4\). The normal area between them is approximately \(0.9192-0.1587=0.7605\) using rounded table values; the calculator’s unrounded area is approximately 0.7606. Report the calculator result rounded to four decimal places.

Conclude. Under the stated normal model, the probability that a randomly selected package takes between 37 and 49 seconds to pass through the scanner is about 0.7606.

A measurement recorded to a chosen precision is still a measurement. For instance, recording scanner times to the nearest whole second makes the recorded values look discrete, but the underlying time is measured on a continuous scale. Match the model to what the variable represents, not only to how a device displays it.

Worked Example: A Sample Proportion and a Normal Approximation

Worked Example: A Sample Proportion and a Normal Approximation

A fictional production line checks 100 items. For each item, the outcome is either acceptable or not acceptable; the process model specifies independent outcomes and probability 0.50 of an item being acceptable. Let \(X\) be the number of acceptable items and let \(\hat{p}\) be the proportion that are acceptable. Approximate the probability that \(\hat{p}\) is at least 0.60.

State. \(X\) counts acceptable items, while \(\hat{p}=X/100\) is their sample proportion. Thus \(\hat{p}\) can take the values 0, 0.01, 0.02, and so on through 1. It is discrete for this fixed sample size, even though its values are written as decimals.

Plan. The count \(X\) meets the binomial conditions: there are 100 fixed trials, two outcomes per trial, independent outcomes, and the same success probability \(p=0.50\). Therefore \(X\sim B(100,0.50)\). Check the Large Counts condition before using a normal approximation: \(np=100(0.50)=50\geq10\) and \(n(1-p)=100(0.50)=50\geq10\). Both hold. The normal model for \(\hat{p}\) has mean \(p=0.50\) and standard deviation \(\sqrt{p(1-p)/n}=0.05\).

Do. The event \(\hat{p}\geq0.60\) is equivalent to \(X\geq60\). Using the continuity correction for the count, approximate this with the normal area above 59.5. On the proportion scale, that boundary is \(59.5/100=0.595\):

$$ P(\hat{p}\geq0.60) =P(X\geq60) \approx\operatorname{normalcdf}(0.595,1E99,0.50,0.05) \approx0.0287 $$

The corresponding standardized boundary is \((0.595-0.50)/0.05=1.9\), whose upper-tail area is approximately 0.0287. This is a normal approximation to the probability, not an exact statement that \(\hat{p}\) has continuous possible values.

Conclude. Under the stated process model, the normal approximation gives a probability of about 0.0287 that at least 60% of the 100 checked items are acceptable.

Common Mistakes and AP Exam Tips

  • Treating every count as binomial. A count is discrete, but it is binomial only when it counts successes in a fixed number of qualifying trials. Full-credit reasoning identifies the trials and checks the conditions rather than relying on the word “count.”
  • Assuming a count over time is automatically binomial. “Number of calls in an hour” gives a count, but does not by itself describe a fixed number of binary trials. A direct probability table may be the appropriate model if probabilities for possible counts are supplied.
  • Calling a sample proportion continuous. For fixed \(n\), \(\hat{p}\) can take only multiples of \(1/n\). A normal model can approximate its distribution under the Large Counts condition, but the actual proportion remains discrete.
  • Assuming every measurement is normal. A measurement is continuous in principle, but that fact alone does not justify a normal model. State that the model is given or refer to evidence about the distribution’s shape.
  • Forgetting what a normal approximation represents. If the original variable is a count or sample proportion, say that the normal model approximates its distribution. For a binomial approximation, check both \(np\geq10\) and \(n(1-p)\geq10\).

A complete AP response makes the connection visible: define the variable in context, describe its possible values, identify the model, and give the information that supports the choice. If a necessary feature is not stated, do not claim it has been verified; explain what information would be needed.

AP Exam Tip: Separate the variable’s type from the model choice. “Discrete count” describes possible values; “binomial” also makes claims about the process. “Measurement” suggests a continuous model, but normality still needs support.

Key Takeaway

Counts, measurements, and proportions point toward different modeling questions. A fixed-trial success count may be binomial; another discrete variable can be modeled with a probability table; and a measurement may fit a normal model when its shape or the stated assumptions support it. A sample proportion has discrete values for fixed \(n\), but its distribution can be approximated by a normal model when the Large Counts condition holds.

Key takeaway: Define what the random variable measures, identify its possible values, and check how those values are generated. Choose a model that fits both the variable type and the process, and label a normal approximation as an approximation.

Check Your Understanding

For each situation, identify the variable type and the model that is supported by the information. Give the key reason for your choice.

  1. A wildlife team counts how many of 20 tagged animals return to a feeding area. The scenario states two possible outcomes per animal, independence, and the same return probability. What model fits the count?
  2. A model gives probabilities for 0, 1, 2, or 3 power outages during a week but does not describe a fixed set of binary trials. What model type does the information support?
  3. A device records the distance traveled by a toy car. The distances are described as roughly symmetric and single-peaked. What model might be appropriate, and what evidence supports it?
  4. In a fixed sample of 50 voters, \(\hat{p}\) is the proportion who support a proposal. What values can \(\hat{p}\) take, and why is it discrete?
  5. A binomial count has \(n=40\) and \(p=0.20\). Check the Large Counts condition. Is a normal approximation justified by that condition?