Tutorials › AP Statistics › Approaching Probability and Random Variable Question Sets

Statistical practices and exam synthesis · Tutorial 1007 of 1020

Approaching Probability and Random Variable Question Sets

Organize a context-based probability set by identifying its random variable and matching each question to the probability model or rule it requires.

Intermediate 10 min read

What You'll Learn

  • Separate the story’s information from the event or random variable each question asks about
  • Use overlap and complements to find probabilities without double-counting
  • Calculate and interpret the expected value of a discrete random variable
  • Recognize and calculate binomial probabilities, including “at least” questions
  • Standardize a normal random variable and find probabilities between or beyond values
  • Check multiple-choice options for mismatched events, units, and rounding

Turn a Shared Context Into a Sequence of Small Questions

A multiple-choice set may describe one setting and then ask several different probability questions about it. The setting stays the same, but the mathematical task can change from one question to the next: one may ask about overlapping events, another about a long-run average, and another about a binomial or normal probability. The most useful first move is to identify what each question asks for, rather than trying to apply one formula to the entire story.

This tutorial brings together probability rules, expected value, binomial models, and normal models. The individual ideas were introduced earlier in the course. Here, the new skill is organizing them efficiently across a set: translate the wording into an event or random variable, match it to the right tool, calculate, and check that the answer makes sense in context.

Key idea: For each question, identify the outcome or random variable, record the information that applies to it, and then choose the rule or model. Do not assume that two questions in the same context require the same calculation.

A Quick Sorting Routine

Read the questions before calculating. Words such as “both,” “either,” and “neither” usually point to event relationships. “Average amount” may ask for expected value. “Exactly,” “at least,” or “no more than” may describe a count. “Between,” “below,” or “above” a value may call for a normal probability if the variable is modeled as normal.

1
Name the target.
Write the event in words or define the random variable. For example, let \(X\) be the number of delayed buses among eight scheduled buses.
2
Match the structure.
For overlapping events, use addition and subtraction of the overlap. For a discrete random variable’s long-run average, use expected value. For a fixed number of independent success-or-failure trials, consider a binomial model. For a normally distributed quantitative variable, standardize to a \(z\)-score.
3
Translate the wording precisely.
“At least one” includes one and every larger count; “exactly two” includes only two. “Neither” is the complement of “at least one” of the named events.
4
Check the result.
Make sure a probability is between 0 and 1, an expected value has the correct units, and the result answers the event actually asked about.

Probability Rules: Keep Track of Overlap

Suppose \(A\) and \(B\) are events. The probability that at least one occurs is found by adding their probabilities and subtracting the probability that both occur. The subtraction matters because outcomes in the overlap would otherwise be counted twice. The complement rule is often a shorter route to “neither” or “at least one.”

Formula: \(P(A\text{ or }B)=P(A)+P(B)-P(A\text{ and }B)\). Also, \(P(\text{not }A)=1-P(A)\). Use “or” inclusively unless the question explicitly says the events cannot happen together.

If the information is given in a table, first identify which counts belong to each event and which count belongs to both. For a randomly selected individual, divide the relevant count by the total number of individuals. This is a practical way to avoid treating “or” as though it meant “exactly one.”

Expected Value: A Long-Run Average, Not a Promise

A discrete random variable has a set of possible numerical values, each with a probability. Its expected value is the sum of each value multiplied by its probability. It describes the long-run average outcome over many repetitions of the same chance process; it does not predict what must happen on one trial.

$$ E(X)=\sum xP(X=x) $$

Check that the probabilities for all possible values add to 1. Expected value uses the values of the random variable, not the probability of just one selected outcome. If the values represent dollars, the expected value is in dollars; if they represent a count, it is in count units.

Binomial and Normal Questions: Recognize the Model First

A binomial random variable counts successes in a fixed number of trials. Before using a binomial calculation, check the conditions: there is a fixed number \(n\) of trials; each trial has two outcomes, success or failure; trials are independent; and the probability \(p\) of success is the same on every trial. If those conditions fit, the probability of exactly \(k\) successes is

$$ P(X=k)=\binom{n}{k}p^k(1-p)^{n-k} $$

For “at least one,” the complement is often simpler than adding several binomial probabilities: subtract the probability of zero successes from 1. The mean number of successes is \(np\). Use the mean to describe the distribution’s long-run center, not as a substitute for a requested probability.

For a normally distributed random variable \(X\) with mean \(\mu\) and standard deviation \(\sigma\), convert a value \(x\) to a standardized value using \(z=(x-\mu)/\sigma\). A normal probability is an area under the model; use a cumulative normal calculation or standard normal table to find it. For a probability between two values, subtract the lower cumulative area from the upper one.

Worked Examples: A Fictional Bus Service

Worked Example: Overlapping Rider Events and a Service Credit

Situation. In a fictional transit survey of 200 riders, 120 say they use the service’s trip-planning app, 90 say they use a monthly pass, and 55 say they use both. Separately, the service offers a randomly selected rider a credit according to this distribution: $0 with probability 0.85, $5 with probability 0.12, and $20 with probability 0.03.

Find the probability a randomly selected surveyed rider uses the app or the pass. Let \(A\) be the event that a rider uses the app and \(B\) the event that the rider uses a pass. The counts give \(P(A)=120/200=0.60\), \(P(B)=90/200=0.45\), and \(P(A\text{ and }B)=55/200=0.275\). Therefore,

$$ P(A\text{ or }B)=0.60+0.45-0.275=0.775 $$

The probability is 0.775, or 77.5%. Subtracting the overlap prevents riders who use both from being counted twice. As a check, the number who use at least one is \(120+90-55=155\), and \(155/200=0.775\).

Find the probability a rider uses neither. “Neither” is the complement of “app or pass,” so \(1-0.775=0.225\). Checking with counts, \(200-155=45\), and \(45/200=0.225\). Both approaches agree.

Find the expected service credit. Let \(C\) be the credit in dollars. First, the probabilities add to \(0.85+0.12+0.03=1\). Then,

$$ E(C)=0(0.85)+5(0.12)+20(0.03)=0+0.60+0.60=\$1.20 $$

Over many repetitions of this credit offer, the average credit per selected rider would approach $1.20. This does not mean that a particular rider is likely to receive exactly $1.20; the possible individual credits are $0, $5, and $20.

Worked Example: Counting Delayed Buses

Situation. For a fictional bus route, suppose each scheduled bus has probability 0.15 of being delayed. Consider eight scheduled buses. Let \(X\) be the number of those buses that are delayed. For this model, assume the eight delay outcomes are independent and the delay probability is the same for every bus.

Check the binomial conditions. There are a fixed \(n=8\) buses; each bus is delayed or not delayed; the outcomes are assumed independent; and each has the same probability \(p=0.15\) of being delayed. Thus \(X\) can be modeled as binomial with \(n=8\) and \(p=0.15\).

Find the probability exactly two buses are delayed. The number of ways to choose two delayed buses from eight is \(\binom{8}{2}=28\). Therefore,

$$ P(X=2)=\binom{8}{2}(0.15)^2(0.85)^6 =28(0.0225)(0.3771495156) \approx 0.2376 $$

The probability is about 0.2376, or 23.76%. The calculation can be checked by multiplying \(0.0225\) by \(0.3771495156\), which gives approximately \(0.0084858641\), and then multiplying by 28 to get approximately \(0.2376042\).

Find the probability at least one bus is delayed. The complement of at least one delay is no delays. By independence, the probability of no delayed buses is \(0.85^8\). Thus,

$$ P(X\geq 1)=1-P(X=0)=1-(0.85)^8 =1-0.2724905250 \approx 0.7275 $$

The chance of at least one delay is approximately 0.7275. As a check using addition, summing \(P(X=k)\) for \(k=1\) through 8 gives the same result, though the complement is more efficient.

Find the expected number of delayed buses. For a binomial random variable, \(E(X)=np\), so \(E(X)=8(0.15)=1.2\) delayed buses. The result is a long-run average across many groups of eight buses, not a claim that 1.2 buses will be delayed in one group.

Worked Example: Wait Times Under a Normal Model

Situation. Suppose individual rider wait times on a fictional route are modeled by a normal distribution with mean \(\mu=10.5\) minutes and standard deviation \(\sigma=1.2\) minutes. Let \(W\) be the wait time for a randomly selected rider.

Find the probability a rider waits between 9 and 12 minutes. Standardize both endpoints:

$$ z_1=\frac{9-10.5}{1.2}=-1.25, \qquad z_2=\frac{12-10.5}{1.2}=1.25 $$

The desired probability is the area between these \(z\)-scores. Using cumulative standard normal values,

$$ P(9<W<12)=P(-1.25<Z<1.25) =\Phi(1.25)-\Phi(-1.25) $$

Using unrounded cumulative values, \(\Phi(1.25)\approx0.8943502\) and \(\Phi(-1.25)\approx0.1056498\). Therefore,

$$ 0.8943502-0.1056498=0.7887004\approx0.7887 $$

About 78.87% of riders are predicted by this model to wait between 9 and 12 minutes. The subtraction can be checked using the symmetry of the normal curve: the two endpoint areas outside this interval are equal, each about 0.1056498, so the area between is \(1-2(0.1056498)=0.7887004\).

Find the probability a rider waits more than 12 minutes. The value 12 minutes is 1.25 standard deviations above the mean. The area to its right is \(1-\Phi(1.25)=1-0.8943502=0.1056498\), or about 0.1056. This is the upper-tail probability, not the area between the mean and 12 minutes.

Multiple-Choice Checks and Common Mistakes

Once a calculation is complete, use the answer choices as a diagnostic. A distractor often reflects a recognizable mistake: adding two event probabilities without subtracting their overlap, treating “at least one” as “exactly one,” or reporting a cumulative normal area when the question asks for an interval. Check the wording and the event before selecting a numerical option.

  • Double-counting an overlap. For “A or B,” include the overlap only once. Subtract \(P(A\text{ and }B)\) after adding the two individual probabilities.
  • Confusing expected value with a guaranteed result. Expected value is a long-run average. It need not be one of the possible outcomes of a single trial.
  • Using a binomial formula without checking conditions. A count is not automatically binomial. Confirm fixed \(n\), two outcomes, independence, and constant \(p\) before using the model.
  • Misreading “at least.” “At least one” includes every positive count. The complement is zero, not exactly one.
  • Using the wrong normal area. A calculator’s cumulative normal result from negative infinity to an endpoint is not automatically the requested interval or upper tail. Subtract or take a complement as needed.
  • Rounding too early or showing inconsistent arithmetic. Keep extra digits during intermediate steps, then round the final probability consistently. If rounded values do not subtract to the displayed result, show more digits or explain that the final value comes from unrounded values.
  • Leaving out units or context. A probability is unitless; an expected number of delayed buses is measured in buses, and an expected credit is measured in dollars. State what the result means for the situation.

A strong response is specific: “The probability that at least one of the eight buses is delayed is about 0.7275 under the stated independent, equal-probability model.” That wording names the event, reports the result, and indicates the assumptions behind the calculation.

Key takeaway: In a context-based question set, classify each task before calculating. Match event wording to probability rules, a distribution of outcomes to expected value, a qualifying count to a binomial model, and a normally modeled measurement to a normal-area calculation. Then check the event, assumptions, arithmetic, and meaning of the answer.

Check Your Understanding

For each question, identify the relevant rule or model before calculating or explaining.

  1. In a group of 150 riders, 80 use a trip-planning app, 60 use a monthly pass, and 30 use both. What is the probability a randomly selected rider uses the app or pass?
  2. A prize is $0 with probability 0.70, $4 with probability 0.20, and $10 with probability 0.10. Find its expected value and explain what it means.
  3. For a binomial variable with \(n=6\) and \(p=0.20\), describe the complement event that makes finding \(P(X\geq1)\) efficient.
  4. What four conditions should be checked before treating a count of delayed buses as binomial?
  5. A normally distributed wait time has mean 8 minutes and standard deviation 2 minutes. What \(z\)-score corresponds to a wait of 12 minutes, and which tail would you use to find the probability of waiting more than 12 minutes?