Tutorials › AP Statistics › The Mean of a Discrete Random Variable

Expected value and variability · Tutorial 321 of 1000

The Mean of a Discrete Random Variable

Use each possible value and its probability to calculate and interpret the mean number of pets in a modeled distribution.

Intermediate 8 min read

What You'll Learn

  • Identify the mean of a discrete random variable as its probability-weighted average.
  • Use the formula mu_X = sum of x times P(X = x).
  • Show and add every value-probability product in a number-of-pets distribution.
  • Interpret the mean with the variable’s context and units.
  • Distinguish the mean from the most likely value and from a possible value of X.

A Weighted Center for a Probability Distribution

In Exam Practice on Discrete Random Variable Distributions, you used a valid distribution to describe the possible values of a random variable and their probabilities. Now we can summarize the distribution’s center with one number: its mean. For a discrete random variable, the mean uses every possible value, with more probable values receiving more weight.

Suppose \(X\) is the number of pets in a randomly selected household. If one value occurs with high probability, it should contribute more to the distribution’s center than a value with very low probability. The calculation accounts for that by multiplying each value \(x\) by its probability \(P(X=x)\), then adding those products.

Definition: The mean of a discrete random variable \(X\), written \(\mu_X\), is the probability-weighted average of all possible values of \(X\). Each value is weighted by the probability that \(X\) takes that value.
Formula: If \(X\) has possible values \(x_1,x_2,\ldots,x_k\), then its mean is the sum of each value multiplied by its probability:
$$ \mu_X=\sum xP(X=x) $$

The notation \(\sum xP(X=x)\) means to calculate \(xP(X=x)\) for every possible value \(x\), then add all the products. A table makes the process visible: include a column for each product so that no value or probability is skipped. The result is the mean of the distribution, not the probability of an event.

As in Notation for Random Variables, \(X\) names the random variable, \(x\) is a possible value, and \(\mu_X\) names the mean of the distribution of \(X\). If \(X\) counts pets, its values and its mean are measured in pets. Probabilities are unitless, so multiplying a value measured in pets by a probability still gives a contribution measured in pets.

Calculate the Mean Row by Row

The most dependable method is to copy each possible value and its probability from the distribution, compute their product, and sum the products. The probability of a value acts as its weight. A value with probability 0 contributes nothing; a value with a larger probability generally makes a larger contribution, though both the value and its probability matter.

Before calculating, use the validity ideas from Checking Whether a Probability Distribution Is Valid: the probabilities should be between 0 and 1, and they should sum to 1. If a table is not a valid distribution, its weighted sum should not be treated as the mean of a valid probability model.

1
Identify the variable and its values.
State what \(X\) represents and list every possible value in the distribution.
2
Multiply each value by its probability.
For every row, calculate \(xP(X=x)\) and show the product.
3
Add the products.
The sum is \(\mu_X\). Keep enough precision in the products until the final result.
4
Interpret the mean in context.
Name the quantity and include its units. Do not claim that the mean must be one of the possible values.

Worked Example: Find the Mean Number of Pets

Consider an invented probability model for the number of pets in a randomly selected household in a neighborhood. Let \(X\) be the number of pets in that household. The model gives the following distribution. Find \(\mu_X\), showing every product.

Number of pets, \(x\)\(P(X=x)\)\(xP(X=x)\)
00.20\(0(0.20)=0.00\)
10.35\(1(0.35)=0.35\)
20.30\(2(0.30)=0.60\)
30.15\(3(0.15)=0.45\)

Check the distribution. Each probability is between 0 and 1, and the probabilities total \(0.20+0.35+0.30+0.15=1.00\). The table is a valid distribution for the listed possible values.

Calculate the mean. Add the products in the last column:

$$ \begin{aligned} \mu_X &=\sum xP(X=x)\\ &=0(0.20)+1(0.35)+2(0.30)+3(0.15)\\ &=0.00+0.35+0.60+0.45\\ &=1.40\text{ pets}. \end{aligned} $$

According to this model, the mean number of pets per household is 1.40. The value 1.40 summarizes the center of the distribution; it does not mean that a household can have 1.40 pets. The possible counts in this model are whole numbers.

What the Products Tell You

Each product \(xP(X=x)\) is that value’s contribution to the weighted mean. For example, in the first example, the two-pet value contributes \(2(0.30)=0.60\) pets to the sum. The zero-pet value has probability 0.20, but its product is \(0(0.20)=0\), so it contributes zero to the mean. It still belongs in the distribution; its contribution being zero does not mean its probability is zero.

The mean is not found by adding the possible values and dividing by the number of rows. That would give each listed value equal weight, even when the probabilities differ. Instead, the formula uses probabilities as weights. If the probabilities were all equal, the weighted mean would match the ordinary average of the possible values, but a discrete distribution does not generally assign equal probabilities to its values.

The mean also need not be the most likely value. The most likely value is the one with the greatest individual probability. The mean uses all the values and probabilities together. A distribution can have one value with the largest probability while its mean lies elsewhere—or even between possible values.

Worked Example: The Mean Is Not Necessarily the Most Likely Count

A second invented model describes \(X\), the number of pets in a randomly selected household in a different community. Find the mean and identify the most likely number of pets.

Number of pets, \(x\)\(P(X=x)\)\(xP(X=x)\)
00.30\(0(0.30)=0.00\)
10.25\(1(0.25)=0.25\)
20.20\(2(0.20)=0.40\)
30.15\(3(0.15)=0.45\)
40.10\(4(0.10)=0.40\)

Check the probabilities. All are between 0 and 1, and \(0.30+0.25+0.20+0.15+0.10=1.00\), so this is a valid distribution.

Find the mean. Include the product from every row:

$$ \begin{aligned} \mu_X &=0(0.30)+1(0.25)+2(0.20)+3(0.15)+4(0.10)\\ &=0.00+0.25+0.40+0.45+0.40\\ &=1.50\text{ pets}. \end{aligned} $$

The most likely number is 0 pets, because \(P(X=0)=0.30\) is the largest probability in the table. The mean is 1.50 pets. These answers describe different features: 0 is the single most probable count, while 1.50 is the probability-weighted center of all five possible counts. Since 1.50 is not a possible value of \(X\), it cannot be the pet count for one household in this model.

Compare Contributions, Not Just Probabilities

A low-probability value can still contribute noticeably to the mean if the value itself is large. This is why the product column is useful: looking only at probabilities can hide how much a higher count affects the weighted sum. In a number-of-pets distribution, a value of 5 has a larger value than 1, but its contribution depends on its probability as well.

A careful calculation retains the full distribution and includes every product—even zero products and small products. You can check your arithmetic in two ways: add the displayed products directly, then substitute the values and probabilities into the formula and evaluate the sum. Both methods should agree. The next example makes the contributions of less common, higher counts visible.

Worked Example: Include a Less Common Higher Count

An invented model for another community gives the number of pets \(X\) in a randomly selected household. Find the mean. Then compare the contribution from households with five pets with the contribution from households with one pet.

Number of pets, \(x\)\(P(X=x)\)\(xP(X=x)\)
00.45\(0(0.45)=0.00\)
10.25\(1(0.25)=0.25\)
20.15\(2(0.15)=0.30\)
30.08\(3(0.08)=0.24\)
40.05\(4(0.05)=0.20\)
50.02\(5(0.02)=0.10\)

Check the distribution. The probabilities are all between 0 and 1. Their sum is \(0.45+0.25+0.15+0.08+0.05+0.02=1.00\).

Calculate the mean. Add every product, including the small contribution for five pets:

$$ \begin{aligned} \mu_X &=0(0.45)+1(0.25)+2(0.15)+3(0.08)+4(0.05)+5(0.02)\\ &=0.00+0.25+0.30+0.24+0.20+0.10\\ &=1.09\text{ pets}. \end{aligned} $$

Compare the two contributions. The one-pet contribution is \(1(0.25)=0.25\) pets. The five-pet contribution is \(5(0.02)=0.10\) pets. Although five is a much larger count, its probability is small, so its contribution is smaller than the one-pet contribution. Both rows still matter to the mean.

According to this model, the mean number of pets per household is 1.09. The calculation uses the five-pet possibility even though it has probability only 0.02; omitting that row would understate the model’s mean.

Common Mistakes and AP Exam Tips

  • Taking an unweighted average. Adding the possible counts and dividing by their number ignores their probabilities. A full-credit calculation weights each count by its probability using \(xP(X=x)\).
  • Multiplying the probabilities together. The mean is a sum of value-probability products, not a product of all the probabilities. Show each row’s product, then add the products.
  • Leaving out a value. Include every possible value in the distribution, even when its probability or its product is small. A missing row changes the sum.
  • Confusing the mean with the mode. The mode is the value with the largest probability; the mean is the weighted sum. Name which feature you found instead of treating them as interchangeable.
  • Insisting the mean must be possible. A mean such as 1.40 pets can be valid even if the random variable only takes whole-number values. Interpret it as the distribution’s center, not as a count for one household.
  • Forgetting units or context. State, for example, “The mean number of pets per household is 1.40 pets,” rather than giving only the number 1.40.
  • Using a table that is not a valid distribution. Check that probabilities are within the allowed range and add to 1 before treating the weighted sum as a distribution mean.

For an AP response, make the calculation easy to verify: identify \(X\), write \(\mu_X=\sum xP(X=x)\), show each product, and add them. Finish with a sentence that names the quantity and its units. Do not describe the mean as the most common outcome or as a value that an individual household must have.

Key Takeaway

The mean of a discrete random variable is a probability-weighted center. Each possible value contributes its value multiplied by the probability attached to it. When \(X\) counts pets, the resulting mean is expressed in pets and can fall between the whole-number counts that \(X\) can take.

Key takeaway: Calculate \(\mu_X\) by showing and adding every product \(xP(X=x)\). Use all possible values, check that the distribution is valid, and interpret the result in context with units.

Check Your Understanding

Use the invented distribution for \(X\), the number of pets in a randomly selected household.

Number of pets, \(x\)\(P(X=x)\)
00.25
10.40
20.25
30.10
  1. Check that this is a valid probability distribution.
  2. Write and calculate each product \(xP(X=x)\), then find \(\mu_X\).
  3. Which number of pets is most likely? Explain how this differs from the mean.
  4. Can the mean be a number of pets that is not a possible value of \(X\)? Explain.
  5. Suppose the probability of 3 pets were 0.20 and all other probabilities stayed the same. Would the table still be a valid distribution? Explain why you should check validity before calculating a mean.