Tutorials › AP Statistics › Common Errors with Random Variables and Distributions

Random variables and distributions · Tutorial 318 of 1000

Common Errors with Random Variables and Distributions

Practice reading random-variable notation carefully and use a short audit to find and fix errors in probability tables.

Intermediate 10 min read

What You'll Learn

  • Distinguish a random variable \(X\) from a particular value \(x\).
  • Read \(P(X=x)\) as the probability that the variable takes a specified value.
  • Check that every listed probability is between 0 and 1 and that the probabilities total 1.
  • Look for omitted possible values when a distribution’s probabilities do not total 1.
  • Use counts and the stated set of possible values to locate errors without guessing at a correction.
  • Write probability statements that match the event and its context.

Small Notation Errors Can Change the Meaning

A discrete probability distribution is compact: it lists possible values of a random variable and the probability attached to each value. That compactness makes the table useful, but it also makes small reading or recording errors easy. A capital letter and a lowercase letter may look similar, for example, but they play different roles. A missing row can also make an otherwise reasonable-looking table incomplete.

In What Is a Random Variable and Probabilities of Single Values, you learned to use a capital letter for a random variable and a lowercase letter for one possible value. In Checking Whether a Probability Distribution Is Valid, you learned the two basic validity checks: each probability must be between 0 and 1, inclusive, and the probabilities for all possible values must add to 1. This tutorial focuses on mistakes that can make those ideas easy to misapply—and on a practical way to catch them.

Definition: In \(P(X=x)\), \(X\) names the random variable, \(x\) is one particular value it can take, and \(P(X=x)\) is the probability that the random variable takes that value.

A useful first move is to read notation aloud in context. For example, if \(X\) is the number of notifications a phone receives during one hour, \(P(X=2)\) means “the probability that the number of notifications in the hour is 2.” The number 2 is a value of the variable; it is not itself the random variable. The variable \(X\) describes the quantity that can vary from hour to hour.

A Quick Audit for Distribution Errors

When a table seems confusing or produces an unexpected answer, pause before calculating. Check what the variable represents, which values it can take, and whether each row matches the correct probability. Then check the probability entries and their total. This order helps separate notation problems from arithmetic or data-entry problems.

1
Name the variable and its unit or meaning.
Read the definition of \(X\). Ask what one value of \(X\) would describe in the setting.
2
Identify the possible values.
Check that the table includes every value the variable can take in the stated model. A missing value can hide probability.
3
Match each value to its probability.
For each row, say, “The probability that \(X\) equals this value is this amount.” Confirm that the value and probability have not shifted into the wrong columns.
4
Check the probability bounds and total.
Every probability must be from 0 to 1, inclusive, and the probabilities for all possible values must sum to 1.

This is an error-checking routine, not a replacement for understanding the chance process. A table might pass the numerical checks while still assigning probabilities to the wrong values. Conversely, a table with a total other than 1 is not automatically fixable by changing any one entry: you need information about the possible values or the original counts to determine what went wrong.

Worked Example: Keep the Variable Separate from Its Value

In an invented phone-use model, let \(X\) be the number of notifications received during one randomly selected hour. The probability distribution is:

Notifications, \(x\)\(P(X=x)\)
00.25
10.40
20.25
30.10

Identify the mistake. A student writes, “\(X=0.25\), so the probability of two notifications is 2.” The sentence confuses the variable, a value, and a probability.

Check the table. The variable \(X\) is the number of notifications in an hour. The value \(x=2\) is one possible notification count. The table entry in that row says \(P(X=2)=0.25\). In words, the model assigns a 0.25 probability to receiving exactly two notifications in a randomly selected hour.

Verify the distribution. Each entry is between 0 and 1. The total is

$$ 0.25+0.40+0.25+0.10=1.00. $$

Correct the statement. The correct notation is \(P(X=2)=0.25\), not \(X=0.25\) or “the probability is 2.” The value 2 counts notifications; 0.25 is a probability.

When Probabilities Do Not Add to 1

A total below 1 can signal that a possible value was left out or that one or more probabilities were recorded incorrectly. A total above 1 can signal an incorrect entry, a duplicated category, or another error. The total is a useful alarm, but it does not by itself diagnose the cause. Start by checking that the rows cover all possible values, then compare the entries with their source.

For a count variable, the possible values often follow a clear sequence, such as 0, 1, 2, 3, and so on up to a stated maximum. If the table jumps from 2 to 4, ask whether 3 is possible and has been omitted. But do not assume that every gap means an error: a variable may have restricted possible values, so its definition and the chance process matter.

Worked Example: Find an Omitted Value Before Changing Probabilities

A school technology team makes an invented model of the number \(X\) of device alerts during a one-hour practice session. The model says \(X\) can be 0, 1, 2, 3, or 4. A draft table lists:

Alerts, \(x\)Draft \(P(X=x)\)
00.20
10.35
20.30
40.10

State the issue. Check whether the table covers the stated values and whether the listed probabilities total 1.

Do the audit. The stated possible values are 0 through 4, but the table omits \(x=3\). The listed probabilities total

$$ 0.20+0.35+0.30+0.10=0.95. $$

The draft is short by \(1-0.95=0.05\). The team checks its original count record and confirms that 5 out of 100 sessions had 3 alerts. Thus, \(P(X=3)=5/100=0.05\), and the missing row supplies exactly the remaining probability.

Verify the correction. The completed probabilities sum to

$$ 0.20+0.35+0.30+0.05+0.10=1.00. $$

Conclude. The corrected table includes all five stated possible alert counts, and its probabilities total 1. The sum alone would have shown that something was wrong, but the stated values and original records identified the missing row and its probability.

A Total Check Does Not Identify Every Wrong Entry

A common overcorrection is to notice that probabilities do not total 1 and then adjust an entry until they do. That can produce a table with a total of 1 but no connection to the chance process. Another mistake is to multiply every listed probability by a constant to force the total to 1. Unless the table is explicitly meant to describe probabilities conditional on the listed outcomes, this changes the model rather than correcting its documented error.

Instead, use the source information when it is available. If probabilities come from relative frequencies, divide each category’s count by the total number of observations. As in Random Variables Defined from Tables of Data, the probability assigned to a value is its count divided by the total count. Then check that the counts cover all observations and that the resulting probabilities sum to 1, allowing for small rounding differences if rounded decimals are shown.

Worked Example: Use Counts to Locate Incorrect Entries

A community garden records the number \(Y\) of seedlings that sprout in a small test tray. In an invented set of 50 trays, the counts are 10 trays with 0 sprouts, 22 with 1, 15 with 2, and 3 with 3. A draft distribution instead lists probabilities 0.20, 0.45, 0.30, and 0.13 in that order.

Check the draft. All four probabilities are individually between 0 and 1, but their total is

$$ 0.20+0.45+0.30+0.13=1.08. $$

So the draft cannot be a valid probability distribution. The total flags a problem, but it does not say which entry is wrong. Compare each entry with the observed count divided by 50:

$$ \begin{aligned} P(Y=0)&=\frac{10}{50}=0.20,\\ P(Y=1)&=\frac{22}{50}=0.44,\\ P(Y=2)&=\frac{15}{50}=0.30,\\ P(Y=3)&=\frac{3}{50}=0.06. \end{aligned} $$

The draft’s entries for 1 and 3 sprouts do not match the counts. The corrected probabilities total

$$ 0.20+0.44+0.30+0.06=1.00. $$

Conclude. The count information identifies the two incorrect entries: the probabilities for 1 and 3 sprouts should be 0.44 and 0.06. Merely subtracting 0.08 from one draft entry could make the sum equal 1, but it would not establish which value or values were recorded incorrectly.

Write the Event You Mean

Another frequent error is to write a value where a probability is requested, or to write a probability without naming the event it describes. For example, if \(X\) counts notifications, “2 notifications” is a possible value, while \(P(X=2)\) is a probability. A statement such as \(P(X)=0.25\) does not specify which event is being assigned that probability.

Before reading a table or adding entries, translate the question into an event using the variable. Then match the event to the appropriate values in the table. The tutorials Cumulative Probabilities for Discrete Variables and Probabilities of Ranges and Inequalities develop how to select values for cumulative and range events. Here, the key error check is to make sure the probability notation names the same event that the words describe.

Worked Example: Match a Probability Statement to the Question

Let \(N\) be the number of new messages received by a help desk during a randomly selected 30-minute period. An invented model gives this distribution:

Messages, \(n\)\(P(N=n)\)
00.15
10.30
20.35
30.20

Question. What is the probability of exactly two messages?

Translate and calculate. “Exactly two” is the event \(N=2\). Read the row for \(n=2\), which gives \(P(N=2)=0.35\). The answer is a probability, not the count 2.

Check the wording. The statement “\(P(N=2)=0.35\)” names both the event and its probability. Saying “\(N=0.35\)” would instead claim that the number of messages is 0.35, which is not what the table describes. Saying “the probability is 2” would confuse the count with the chance.

Interpret. According to this model, a randomly selected 30-minute period has probability 0.35 of containing exactly two new messages.

Common Mistakes and AP Exam Tips

  • Using \(X\) and \(x\) as if they were interchangeable. \(X\) is the random variable; \(x\) is a particular possible value. Write \(P(X=x)\) when describing the probability of that value.
  • Giving a value instead of a probability. If asked for \(P(X=2)\), report the probability from the row for 2, not the value 2 itself. Interpret the result using what \(X\) counts or measures.
  • Checking only whether each entry is between 0 and 1. That check is necessary but not sufficient. Add the probabilities for the complete set of possible values as well.
  • Assuming a total below 1 tells you exactly what is missing. It might indicate an omitted value, but it could also reflect an incorrect entry. Check the stated possible values and the source information.
  • Changing an entry just to force a total of 1. A total of 1 does not prove that each probability is correct. Use the original counts or model to verify the individual entries.
  • Leaving the event unclear. A full-credit response makes clear what \(X=x\) means in context. For example: “The probability that a randomly selected hour has exactly two notifications is 0.25.”

For a clear AP response, define the variable or refer to its stated definition, write the probability event precisely, and give the probability attached to that event. When auditing a table, state which check fails and use the available information to support any correction. Do not claim a particular entry is wrong based only on the total unless the other information identifies it.

Key Takeaway

Distribution errors are easier to catch when you keep the variable, its possible values, and their probabilities distinct. First check the meaning and coverage of the table; then check the individual probability bounds and the total.

Key takeaway: \(X\) names the random variable, \(x\) is a possible value, and \(P(X=x)\) is the probability of that value. A valid discrete distribution includes all possible values, assigns each a probability from 0 to 1, inclusive, and has probabilities that sum to 1.

Check Your Understanding

Use the invented distribution below, where \(K\) is the number of packages delivered to a particular building during a randomly selected hour.

Packages, \(k\)\(P(K=k)\)
00.10
10.25
20.40
30.25
  1. In context, explain what \(K\), \(k=2\), and \(P(K=2)\) each mean.
  2. What is the probability of exactly two packages? Write the notation and interpret the result in context.
  3. Check that every probability is between 0 and 1, inclusive, and that the probabilities sum to 1.
  4. A student writes \(K=0.40\) to report the probability of two packages. Explain the notation error and give the corrected statement.
  5. If the row for \(k=1\) were accidentally omitted, what would the listed probabilities total? What information would you check before deciding how to correct the table?