Tutorials › AP Statistics › Using Relative Frequencies to Estimate Probabilities

Categorical tables and summaries · Tutorial 37 of 1000

Using Relative Frequencies to Estimate Probabilities

Use a table’s relative frequencies to estimate the chance that a randomly chosen individual from the sample belongs to a specified category or group.

Beginner 9 min read

What You'll Learn

  • Explain how a sample’s relative frequencies can be used as empirical probability estimates.
  • Identify the sample space and the event when reading a categorical table.
  • Estimate probabilities for one category, combined categories, and a category’s complement.
  • Distinguish joint, marginal, and conditional probabilities in a two-way table.
  • State what a sample-based probability does and does not say about a wider population.

From Relative Frequency to Estimated Probability

In Frequency Tables and Relative Frequency, you learned to divide each category count by the total number of responses. In Simpson’s Paradox in Categorical Data, you saw why it matters to keep the group and denominator clear. Here, we give relative frequencies a probability interpretation: they can estimate the chance that a randomly chosen individual from the sample has a specified category.

Imagine selecting one individual at random from the people represented in a table, with each individual equally likely to be selected. The sample’s relative frequency for a category tells us what fraction of those individuals belong to that category. We can use that fraction as an estimated probability for the random selection.

Definition: An empirical probability estimate uses observed relative frequency as an estimate of the chance of an event. For a category, divide the number of sampled individuals in that category by the total number of sampled individuals. The result is the estimated probability that a randomly chosen individual from the sample belongs to that category.

For example, if 18 out of 80 sampled individuals are in a category, the relative frequency is \(18/80=0.225\). In a model based on this sample, the estimated probability of selecting an individual in that category is \(0.225\), or \(22.5\%\). The count is not itself the probability; the count divided by the appropriate total is.

The set of individuals represented in the table is the sample space for this random selection. An event is the outcome or group of outcomes we are interested in, such as selecting someone who bikes to school. Be clear about both: “randomly select one student from these 80 students” describes the sample space, while “the student bikes to school” describes the event.

Formula: For an event \(A\), estimate its probability by dividing the number of sampled individuals for whom \(A\) is true by the total sample size \(n\).
$$ \widehat{P}(A)=\frac{\text{number of sampled individuals in event }A}{n} $$
The symbol \(\widehat{P}(A)\) means an estimated probability for event \(A\), based on the sample’s relative frequency.

Using a One-Variable Table

A one-variable frequency table gives the count in each category. Divide a category count by the table’s total to find its relative frequency and estimated probability. If the table already displays relative frequencies, those values can be read as the probability estimates directly, as long as the random selection is from the individuals represented in that table.

When the categories are mutually exclusive and account for every individual, their relative frequencies add to 1. That makes sense: a selected individual must belong to one of the listed categories. If an event includes more than one category, add the counts for those categories and divide by the original total. Equivalently, add their relative frequencies.

Worked Example: How a Student Gets to School

A fictional school survey records the usual way 80 sampled students travel to school. Suppose one student is selected at random from these 80. Estimate the probability that the student walks or bikes.

Usual travel methodFrequencyRelative frequency
Walk220.275
Bike100.125
Bus180.225
Car300.375
Total801.000

The event is “the selected student walks or bikes.” These are separate categories, so add their counts and divide by the total:

$$ \widehat{P}(\text{walk or bike}) =\frac{22+10}{80} =\frac{32}{80} =0.400 $$

The same result comes from adding the two relative frequencies: \(0.275+0.125=0.400\). Thus, based on these sample data, the estimated probability that a randomly chosen student from the 80 sampled students walks or bikes to school is \(0.400\), or \(40\%\).

As a check, the four category frequencies add to \(22+10+18+30=80\), and their relative frequencies add to \(0.275+0.125+0.225+0.375=1.000\). This confirms that the categories account for the full sample.

Events in a Two-Way Table

A two-way table organizes individuals according to categories of two categorical variables. As covered in Joint Relative Frequencies, a cell count divided by the grand total is a joint relative frequency. It can be read as an estimated probability for selecting an individual who meets both category conditions. A marginal relative frequency describes one category across the full sample. A conditional relative frequency uses only the specified group as its denominator, as in the earlier tutorials on conditional distributions.

These are different questions, so they can have different answers. “Select a student who is in Grade 10 and brings a reusable bottle” refers to a joint event and uses the grand total. “Select a student who brings a reusable bottle” refers to a marginal event and also uses the grand total. “Among Grade 10 students, select one who brings a reusable bottle” is conditional on being in Grade 10 and uses the Grade 10 total.

Key distinction: For a joint or marginal event, use the grand total if the selection is from the entire table. For a conditional event, the wording specifies a group to select from, so use that group’s total. State the selection group before choosing the denominator.

Worked Example: Reusable Bottles by Grade

A fictional school survey asks 120 students whether they usually bring a reusable water bottle. The two-way table shows the responses by grade. Estimate three probabilities for a random selection from the 120 surveyed students.

GradeBrings a bottleDoes not bring a bottleTotal
Grade 9241640
Grade 10301040
Grade 11202040
Total7446120

First, estimate the probability that a randomly selected surveyed student is in Grade 10 and brings a bottle. This is a joint event, so use the Grade 10-and-bottle cell over the grand total:

$$ \widehat{P}(\text{Grade 10 and brings a bottle}) =\frac{30}{120} =0.250 $$

Next, estimate the probability that the selected student brings a bottle, regardless of grade. This is a marginal probability. Add the counts in the “brings a bottle” column, or use its column total:

$$ \widehat{P}(\text{brings a bottle}) =\frac{74}{120} \approx0.6167 $$

Finally, estimate the probability that a student brings a bottle given that the selection is from Grade 10 students. The phrase “given that” identifies the group: the 40 Grade 10 students. The Grade 10-and-bottle count is the numerator:

$$ \widehat{P}(\text{brings a bottle}\mid\text{Grade 10}) =\frac{30}{40} =0.750 $$

The answers are different because the questions describe different selection processes. From all 120 surveyed students, the estimated chance of selecting someone who is both in Grade 10 and brings a bottle is \(0.250\). From all 120, the estimated chance of selecting someone who brings a bottle is about \(0.6167\). If the selection is restricted to Grade 10 students, the estimated chance of bringing a bottle is \(0.750\).

What the Estimate Applies To

A relative frequency is an observed summary of a sample. If the random choice is made from the sample itself, its relative frequency gives the exact proportion of the sample in the event and serves as the probability estimate for that choice. For instance, exactly 74 of the 120 students in the bottle example bring one, so the share of those surveyed students who would meet that description is \(74/120\).

Using the sample relative frequency to say something about a larger population is a different step. The sample-based value may be used as an estimate of a population probability, but how reasonable that estimate is depends on how the sample was obtained and how well it represents the population. A voluntary online poll, for example, may not represent all members of the group the poll is about. A table’s relative frequency alone does not prove that the same proportion holds in the wider population.

Keep the wording matched to the evidence. “In this sample, 74 of 120 students bring a reusable bottle” describes the observed data. “For a random selection from these 120 students, the estimated probability of selecting one who brings a bottle is about \(0.6167\)” describes the sample-based probability. A claim about all students at the school needs a sampling design that supports extending the result beyond those surveyed.

Worked Example: A Voluntary Trail-Use Poll

A fictional recreation group posts an optional online poll asking hikers whether they support adding more trail signs. Of the 100 people who respond, 68 say yes and 32 say no. Estimate the probability of selecting a “yes” respondent at random from the 100 responses, and explain what the estimate does not establish.

The event is selecting a respondent who answered yes. The relevant count is 68 and the sample size is 100:

$$ \widehat{P}(\text{yes among respondents}) =\frac{68}{100} =0.68 $$

Thus, if one of these 100 respondents is selected at random, the estimated probability of selecting a person who answered yes is \(0.68\), or \(68\%\). The 68% is an accurate description of this set of responses.

However, the poll was voluntary. People who chose to respond may differ from hikers who did not see or answer it. Therefore, the table by itself does not establish that 68% of all hikers support adding signs. To make a broader claim, we would need to consider how the respondents were selected and whether they represent the population of interest.

Conclusion: The sample relative frequency estimates the probability of a “yes” answer for a randomly selected poll respondent. It is not automatically a reliable estimate of the opinion of every hiker.

Common Mistakes and AP Exam Tips

  • Using a count as though it were a probability. A frequency such as 30 is a count. Divide by the correct total to obtain the relative frequency, \(30/120=0.250\).
  • Choosing the wrong denominator. The denominator depends on the selection described. For a random selection from everyone in a two-way table, use the grand total; for a selection among one row or column group, use that group’s total.
  • Confusing a joint event with a conditional event. “Grade 10 and brings a bottle” uses the cell count over the grand total. “Brings a bottle among Grade 10 students” uses the same cell count over the Grade 10 total.
  • Adding categories that do not match the event. To find the probability of walking or biking, include those two categories—not the entire table or categories outside the event.
  • Claiming a sample proportion is automatically a population fact. Describe the sample-based estimate first. Extend it to a population only when the way the sample was selected supports that interpretation.

For a clear answer, identify the random selection group, name the event, show the count divided by the proper total, and interpret the result in context. When the question says “given that” or “among,” state the restricted group before calculating. Keep the conclusion limited to the sample unless the sampling method justifies a broader claim.

Key takeaway: A table’s relative frequency can be used as an estimated probability for a randomly selected individual from the sample. Identify the event and selection group, then use the matching count and denominator. A sample-based probability does not automatically describe a wider population.

Check Your Understanding

For each question, identify the event and use a denominator that matches the stated selection group.

  1. A survey of 50 students records 12 who walk, 8 who bike, 20 who take a bus, and 10 who travel by car. Estimate the probability that a randomly selected surveyed student bikes.
  2. Using the same survey, estimate the probability that a randomly selected surveyed student walks or bikes. Show the calculation using counts and using relative frequencies.
  3. In a two-way table, a cell contains 18 individuals out of a grand total of 90. What kind of event does the cell’s relative frequency describe? Calculate it.
  4. A table shows 18 individuals in a specified cell and a row total of 30. What conditional probability could \(18/30\) estimate? Explain why the denominator is not the grand total.
  5. A voluntary poll finds that 72% of respondents favor a proposal. What probability does this relative frequency estimate for a random selection from the respondents, and why might it not estimate the opinion of the entire population well?