Tutorials › AP Statistics › Expected Value from a Table of Relative Frequencies

Expected value and variability · Tutorial 330 of 1000

Expected Value from a Table of Relative Frequencies

Use relative frequencies as empirical probabilities to find and interpret the expected count for one observational unit.

Intermediate 9 min read

What You'll Learn

  • Convert counts in a frequency table into relative-frequency probabilities
  • Calculate an empirical expected count using a weighted sum
  • Use the total of value-times-frequency products divided by the sample size
  • Explain why the empirical expected count equals the observed sample mean
  • Interpret the result as a long-run model average, not a guaranteed outcome
  • Check a calculation with 1-Var Stats using frequencies as weights

From a Frequency Table to an Expected Count

In Finding Mean and Standard Deviation with 1-Var Stats, you learned that a probability list can be used as a frequency list to calculate summaries of a discrete random variable. Here, the table begins with observed frequencies rather than probabilities. You will use each frequency’s share of the total as an empirical probability, then calculate the expected count.

Suppose a table records the number of items needing repair in each inspected batch. A frequency of 12 for \(X=2\) means 12 batches had exactly 2 items needing repair. If the table summarizes 40 batches, the relative frequency is \(12/40=0.30\). Treating that relative frequency as an empirical probability says that, according to this model, a randomly selected comparable batch has a 0.30 probability of having exactly 2 items needing repair.

Definition: An empirical probability distribution assigns each value \(x\) of a discrete random variable \(X\) a probability equal to its observed relative frequency. If \(f_x\) observations have value \(x\) among \(n\) observations, then \(P(X=x)=f_x/n\) in the empirical model.

Once the relative frequencies are in place, the expected value is calculated as in The Mean of a Discrete Random Variable: multiply each possible value by its probability and add the products. Because each probability is \(f_x/n\), the calculation can also be made directly from the frequency table.

$$ \mu_X =\sum xP(X=x) =\sum x\left(\frac{f_x}{n}\right) =\frac{\sum xf_x}{n}, \qquad n=\sum f_x. $$

The last form is often the quickest. Multiply each count value by how many times it occurred, add those products, and divide by the total number of observations. The result is the observed sample mean, \(\bar{x}\). Equivalently, it is the mean of the probability distribution formed by treating the observed relative frequencies as probabilities.

Key distinction: A frequency \(f_x\) is a number of observations; a relative frequency \(f_x/n\) is a proportion and serves as the empirical probability. To find the mean, either use the relative frequencies as probabilities or divide the total of the value-times-frequency products by \(n\).

A Reliable Calculation Process

A frequency table gives the information needed for an empirical expected value, but the table must represent the observations consistently. Identify what one observation is—for example, one day, one order, or one inspected batch—and what \(X\) counts for that observation. Then make sure the listed frequencies account for all observations.

1
Define \(X\) and identify the observational unit.
State what is counted and the period or unit to which each count refers.
2
Find the total number of observations.
Add the frequencies to get \(n\). This is the denominator for every relative frequency.
3
Calculate the empirical probabilities.
For each value \(x\), divide its frequency \(f_x\) by \(n\). The relative frequencies should total 1, apart from minor rounding.
4
Find the probability-weighted mean.
Calculate \(xP(X=x)\) for each row and add the products, or calculate \(\sum xf_x/n\).
5
Interpret the result in context.
Include what is being counted and the unit of observation. If using the empirical model to discuss future observations, make clear that the mean describes an average over many comparable repetitions, not a promised result for one observation.

The empirical distribution is a way to use observed proportions as a probability model, as introduced in Random Variables Defined from Tables of Data. It does not make the observed pattern certain to repeat exactly. It uses the data’s relative frequencies as the probabilities for the model.

Worked Example: Daily Support Requests

Worked Example: Daily Support Requests

Suppose a fictional support team records the number of requests arriving during each of 40 workdays. Let \(X\) be the number of requests during one workday. Use the frequency table to find the expected count in the empirical probability model.

Requests \(x\)Frequency \(f_x\)Relative frequency \(f_x/40\)\(xP(X=x)\)
060.150.00
1140.350.35
2120.300.60
360.150.45
420.050.20
Total401.001.60

Check the total and calculate the probabilities. The frequencies add to \(6+14+12+6+2=40\), so \(n=40\). Dividing each frequency by 40 gives the relative frequencies in the table. They add to \(0.15+0.35+0.30+0.15+0.05=1.00\), as required for a probability distribution.

Calculate the expected count. Using the relative frequencies, multiply each request count by its probability and add:

$$ \begin{aligned} \mu_X &=0(0.15)+1(0.35)+2(0.30)+3(0.15)+4(0.05)\\ &=0+0.35+0.60+0.45+0.20\\ &=1.60\text{ requests per workday}. \end{aligned} $$

Check the calculation directly from the frequencies. The total of the value-times-frequency products is \(0(6)+1(14)+2(12)+3(6)+4(2)=0+14+24+18+8=64\). Dividing by the 40 workdays gives \(64/40=1.60\) requests per workday, the same result.

The empirical expected count is also the average number of requests across these 40 observed workdays. If the relative-frequency model is used for comparable future workdays, its expected value is 1.60 requests per workday. This does not mean that a particular workday must have a fractional number of requests or exactly 1.60 requests; it is a probability-weighted average.

Worked Example: Late Buses

Worked Example: Late Buses

In a fictional record of 50 school days, a transportation coordinator counts how many buses arrive late on each day. Let \(L\) be the number of late buses on one day. Find the empirical probability of exactly two late buses and the expected number of late buses per day.

Late buses \(l\)Days observed
028
112
27
33
Total50

Find the empirical probability for two late buses. Seven of the 50 days had exactly two late buses, so the relative frequency used as the empirical probability is \(P(L=2)=7/50=0.14\). In this model, the probability of exactly two late buses on a randomly selected comparable day is 0.14.

Calculate the mean count. The relative frequencies for \(L=0,1,2,3\) are \(28/50=0.56\), \(12/50=0.24\), \(7/50=0.14\), and \(3/50=0.06\). They sum to 1.00. Weight each count by its relative frequency:

$$ \begin{aligned} \mu_L &=0(0.56)+1(0.24)+2(0.14)+3(0.06)\\ &=0+0.24+0.28+0.18\\ &=0.70\text{ late buses per day}. \end{aligned} $$

The direct frequency calculation confirms the result. The weighted total is \(0(28)+1(12)+2(7)+3(3)=0+12+14+9=35\) late buses across 50 days. Therefore, \(35/50=0.70\) late buses per day. In context, the observed average was 0.70 late bus per day, and the empirical model’s expected count is 0.70 late buses per day.

Notice that the value \(L=0\) is included. Days with no late buses are still observations in the data and contribute to the relative-frequency model, even though their value-times-frequency product is zero.

Worked Example: Phone Notifications

Worked Example: Phone Notifications

Suppose a fictional class records the number of phone notifications each student receives during a particular 30-minute study period. The table summarizes 80 students. Let \(N\) be the number of notifications received by one student during that period. Find the expected count and describe a calculator check.

Notifications \(n\)Frequency \(f_n\)Relative frequency\(nf_n\)
080.100
1200.2520
2280.3556
3160.2048
480.1032
Total801.00156

Use the relative frequencies as probabilities. The frequencies total \(8+20+28+16+8=80\). The relative frequencies are \(8/80=0.10\), \(20/80=0.25\), \(28/80=0.35\), \(16/80=0.20\), and \(8/80=0.10\). Their sum is 1.00. The probability-weighted mean is:

$$ \begin{aligned} \mu_N &=0(0.10)+1(0.25)+2(0.35)+3(0.20)+4(0.10)\\ &=0+0.25+0.70+0.60+0.40\\ &=1.95\text{ notifications per student per 30-minute period}. \end{aligned} $$

As a check, the frequency-weighted total is \(0(8)+1(20)+2(28)+3(16)+4(8)=0+20+56+48+32=156\). Dividing by 80 students gives \(156/80=1.95\) notifications per student per 30-minute period.

Check with 1-Var Stats. Enter 0, 1, 2, 3, and 4 in L1, and enter 8, 20, 28, 16, and 8 in L2. Run 1-Var Stats with L1 as the data list and L2 as the frequency list. The calculator’s \(\bar{x}\) should be 1.95. Using observed frequencies as weights calculates the same weighted mean as using the corresponding relative frequencies, because each relative frequency is the frequency divided by the same total, 80.

In context, the mean number of notifications recorded per student during the specified study period was 1.95. The empirical model uses these observed proportions to represent the distribution of notifications for a comparable student and period.

Common Mistakes and AP Exam Tips

  • Using frequencies as if they were probabilities. A frequency such as 12 is a count of observations, not a probability. Divide it by the total number of observations to obtain its relative frequency, or use the direct formula and divide the weighted total by \(n\).
  • Forgetting to divide by the total. The sum \(\sum xf_x\) is a weighted total, not the mean. Divide it by \(n=\sum f_x\). In the notifications example, 156 is the total number of notifications recorded, while 1.95 is the average count per student.
  • Leaving out observations with a count of zero. A zero count contributes zero to \(\sum xf_x\), but its frequency still belongs in the total \(n\). Excluding those observations would change the denominator and overstate the average.
  • Switching the observational unit. A mean of 0.70 late buses per day is not the mean per bus or per week. Name the unit used to collect each observation when interpreting the result.
  • Calling the expected count a guaranteed outcome. The mean can be a non-integer because it is a weighted average. It describes the empirical model’s center and does not predict the exact count on one occasion.
  • Treating observed relative frequencies as certainty about the future. Relative frequencies define an empirical model from the data. They do not guarantee that future observations will occur in those same proportions.

A clear response shows the total frequency, the relative-frequency calculation or the direct weighted-total calculation, and an interpretation with units. For example: “The expected number of late buses is \(35/50=0.70\) per day. According to the empirical model, the average number of late buses over many comparable days is 0.70 per day; a single day will have a whole-number count.”

Key Takeaway

A frequency table can be treated as an empirical probability distribution by dividing each frequency by the total number of observations. Weight each count by its relative frequency and add, or divide the total of count-times-frequency products by the sample size. The result is the observed mean count and the expected count under the empirical model.

Key takeaway: If \(f_x\) observations have value \(x\) out of \(n\) total observations, use \(P(X=x)=f_x/n\) and calculate \(\mu_X=\sum xP(X=x)=\sum xf_x/n\). Interpret the answer using the count and the observational unit.

Check Your Understanding

Use the fictional table below, which records the number of damaged packages in each of 40 inspected deliveries. Let \(D\) be the number of damaged packages in one delivery.

Damaged packages \(d\)Frequency \(f_d\)
018
114
26
32
  1. Find the total number of deliveries and calculate the empirical probability \(P(D=2)\).
  2. Write the relative frequency for each possible value of \(D\), and check that the probabilities sum to 1.
  3. Calculate \(\mu_D\) using \(\sum dP(D=d)\), showing each product.
  4. Calculate the mean directly using \(\sum df_d/n\). Confirm that it matches your answer to question 3.
  5. Interpret the expected count in context, and explain why it does not mean that one delivery will contain a fractional number of damaged packages.