Tutorials › AP Statistics › Comparing Two Competing Models

Probability model interpretation · Tutorial 393 of 1000

Comparing Two Competing Models

Check each model’s validity, compare its expected outcomes with observed data, and explain what the comparison does—and does not—show.

Intermediate 10 min read

What You'll Learn

  • Check that competing models describe the same mutually exclusive and collectively exhaustive outcomes.
  • Verify that every proposed probability is between 0 and 1 and that each model’s probabilities sum to 1.
  • Calculate expected counts from a model and the number of observed trials.
  • Compare expected counts and model probabilities with observed counts and relative frequencies.
  • Describe which valid model is closer to the data while accounting for chance variation and context.
  • Explain why a closer fit to one data set does not prove that a model is true.

Two Models, One Chance Process

In Revising a Probability Model, you saw how evidence can motivate changing a model’s probabilities or assumptions. Sometimes, however, the evidence does not point to a single obvious revision. Two different models may both be plausible descriptions of the same process. To compare them, first check that each is a valid probability model, then examine how well each model’s predictions match data collected from that process.

The comparison only makes sense when the models describe the same outcomes for the same process. If one model groups outcomes differently, or if the observed data came from different conditions, a direct comparison can be misleading. As in Checking Outcomes and Probabilities in a Model, verify that the outcomes are mutually exclusive and collectively exhaustive before comparing probabilities assigned to them.

Definition: Comparing two competing models means checking whether each model is a valid probability distribution for the process, then describing how closely each model’s predictions agree with relevant observed data. A closer match in one data set is evidence of better fit to that data set; it does not prove that model is true.

Check Validity Before Comparing Fit

For each proposed distribution, check that every probability is between 0 and 1, inclusive, and that the probabilities across all possible outcomes sum to 1. Also confirm that both models assign probabilities to the same outcome categories. If one proposal fails these checks, it is not a valid probability model as stated, so it should not be treated as a serious competitor on the basis of its apparent fit.

Validity and fit are different questions. Validity is an internal check on the probability assignments: do they satisfy the rules for a probability model? Fit is a comparison with data: are the outcomes observed at frequencies reasonably close to what the model predicts? A model can be valid but fit a particular set of data poorly. It can also look close to a data set while being invalid, in which case its probabilities still need correction.

Comparison checklist:
  • Use the same outcome categories for both models and the observed data.
  • Check that the categories are mutually exclusive and collectively exhaustive.
  • For each model, check that every probability is between 0 and 1 and that the probabilities sum to 1.
  • Make sure the data describe the process under conditions relevant to the models.
  • Compare observed relative frequencies or counts with the corresponding model probabilities or expected counts.
  • Describe the comparison cautiously: data can vary from one set of trials to another.

Worked Example: One Proposed Model Is Invalid

A device sorts small parcels into one of four destination bins, labeled A, B, C, and D. Two proposals give the probabilities of a parcel going to each bin. Model 1 assigns probabilities \(0.40, 0.30, 0.20,\) and \(0.10\). Model 2 assigns probabilities \(0.30, 0.30, 0.25,\) and \(0.20\). Check whether both proposals are valid probability models before comparing them with data.

ModelABCDSum
Model 10.400.300.200.101.00
Model 20.300.300.250.201.05

The bins name all possible destinations, and a parcel goes into exactly one bin. The outcome categories therefore fit this comparison. Every listed probability is between 0 and 1, but Model 2’s probabilities sum to more than 1.

$$ \begin{aligned} \text{Model 1: }&0.40+0.30+0.20+0.10=1.00\\ \text{Model 2: }&0.30+0.30+0.25+0.20=1.05 \end{aligned} $$

Conclusion. Model 1 is a valid probability distribution. Model 2 is not valid as stated because its probabilities sum to \(1.05\). Even if Model 2’s individual probabilities look plausible, they cannot all represent the probabilities of these four mutually exclusive, collectively exhaustive outcomes. The probabilities would need to be corrected before Model 2 could be compared as a probability model.

Compare Predictions with Observed Outcomes

For a valid model, compare its probabilities with the observed relative frequencies: the observed count in each category divided by the total number of observations. Probabilities and relative frequencies are on the same scale, so this comparison shows which categories the model makes too common or too rare relative to the data.

You can also compare counts. If there are \(n\) observations and a model assigns probability \(p_i\) to category \(i\), the model’s expected count for that category is \(np_i\). This is the count the model predicts on average over repeated sets of \(n\) observations, not a promise that the actual count will equal it. As explained in Using Data to Check a Probability Model, observed counts naturally differ from expected counts.

Formula: For \(n\) observations and model probability \(p_i\) for category \(i\), the expected count is \(np_i\). Compare that count with the observed count in the same category; alternatively, compare \(p_i\) with the observed relative frequency for that category.

A simple descriptive comparison is the absolute difference between the observed and expected counts in each category. Absolute differences describe the size of the mismatch without allowing positive and negative differences to cancel. Adding them gives a rough summary of total discrepancy for that particular data set. This is an informal comparison tool, not a formal significance test; it does not account for every feature of the probability process, and it should not replace examining the category-by-category pattern.

Worked Example: Comparing Two Valid Parcel Models with Data

A sorting device sends parcels to bins A, B, C, and D. In 120 parcels processed under the conditions of interest, the observed counts are 42, 33, 27, and 18, respectively. Compare Model 1, with probabilities \(0.40, 0.30, 0.20,\) and \(0.10\), with Model 2, with probabilities \(0.36, 0.28, 0.22,\) and \(0.14\).

Both models assign probabilities to the same complete set of bins. Model 1’s probabilities sum to \(1.00\), and Model 2’s sum to \(0.36+0.28+0.22+0.14=1.00\). Every probability is between 0 and 1, so both models are valid. Now calculate each model’s expected counts using \(n=120\), and compare them with the observed counts.

BinObserved countModel 1 expected countModel 2 expected count
A42\(120(0.40)=48\)\(120(0.36)=43.2\)
B33\(120(0.30)=36\)\(120(0.28)=33.6\)
C27\(120(0.20)=24\)\(120(0.22)=26.4\)
D18\(120(0.10)=12\)\(120(0.14)=16.8\)

For Model 1, the absolute count differences are \(6, 3, 3,\) and \(6\), totaling \(18\). For Model 2, the differences are \(1.2, 0.6, 0.6,\) and \(1.2\), totaling \(3.6\). These totals come from adding the category-by-category differences, not from subtracting the sum of the counts: both the observed and expected counts total 120.

$$ \begin{aligned} D_1&=|42-48|+|33-36|+|27-24|+|18-12|=18\\ D_2&=|42-43.2|+|33-33.6|+|27-26.4|+|18-16.8|=3.6 \end{aligned} $$

The observed relative frequencies are \(42/120=0.350\), \(33/120=0.275\), \(27/120=0.225\), and \(18/120=0.150\). These are close to Model 2’s probabilities in all four bins and are farther from Model 1’s probabilities. Model 2 therefore fits these observed data better by this descriptive comparison.

Conclusion. Both proposed distributions are valid, but Model 2’s expected counts are closer to the observed parcel counts in this set of 120. This is evidence that Model 2 describes these data better than Model 1 does. It does not establish that Model 2 is the true distribution or guarantee that it will fit future data better.

Look at the Pattern, Not Just One Summary

A single total discrepancy can hide useful information. Check which categories have more observations than a model predicts and which have fewer. A model might be close overall but consistently miss one important category. Also check whether its probabilities are close to the observed relative frequencies across the full outcome list, rather than focusing only on the largest count or one event of interest.

The size of a discrepancy should be interpreted in relation to the number of observations. A difference of two parcels may be notable in a set of ten but small in a set of several hundred. Looking at relative frequencies helps make the scale explicit: compare each observed count divided by \(n\) with the probability assigned by the model. In either form, remember that counts from repeated trials can vary even when a model is reasonable.

Data quality matters as well. The comparison is informative only if the recorded outcomes are accurate and the observations represent the process the models are meant to describe. If a machine was adjusted midway through data collection, for example, combining all observations could conceal a change in the process. As discussed in Limitations of Probability Models, a mismatch may come from the model, the data, or the conditions under which the data were collected.

Worked Example: Rechecking the Models with a Smaller Batch

The same parcel-sorting process is monitored during a separate batch of 60 parcels. The observed counts in bins A, B, C, and D are 23, 16, 13, and 8. Compare the same two valid models from the previous example: Model 1 assigns probabilities \(0.40, 0.30, 0.20,\) and \(0.10\); Model 2 assigns probabilities \(0.36, 0.28, 0.22,\) and \(0.14\).

The counts sum to \(23+16+13+8=60\). Multiply each model probability by 60 to get its expected counts. Then compare the predictions with the observed counts by category.

BinObserved countModel 1 expected countModel 2 expected count
A23\(60(0.40)=24\)\(60(0.36)=21.6\)
B16\(60(0.30)=18\)\(60(0.28)=16.8\)
C13\(60(0.20)=12\)\(60(0.22)=13.2\)
D8\(60(0.10)=6\)\(60(0.14)=8.4\)
$$ \begin{aligned} D_1&=|23-24|+|16-18|+|13-12|+|8-6|=6\\ D_2&=|23-21.6|+|16-16.8|+|13-13.2|+|8-8.4|=2.8 \end{aligned} $$

Model 2 again has the smaller total absolute difference. The observed relative frequencies are \(23/60\approx0.3833\), \(16/60\approx0.2667\), \(13/60\approx0.2167\), and \(8/60\approx0.1333\). Their pattern is closer to Model 2’s probabilities than to Model 1’s probabilities.

Conclusion. For this separate batch of 60 parcels, Model 2 is descriptively closer to the observed counts. The comparison points in the same direction as the 120-parcel batch, but it still does not prove that Model 2 is correct. Each batch is a limited set of outcomes, and the process could vary over time.

What a Model Comparison Can Establish

A careful comparison supports a limited conclusion: among the valid proposals considered, one may be closer to the observed outcomes in the data examined. It does not show that all other possible models are wrong, that the closer model is exactly correct, or that a future set of outcomes will match its predictions. A model’s usefulness depends on how well its assumptions and probabilities represent the process of interest, not just on one numerical match.

It is also possible for the better-fitting model to change when new data are collected. If two models are similarly close, say so rather than overstating a small difference. If one is clearly closer across the categories and in additional relevant data, describe that as stronger descriptive support, while still recognizing that observations vary.

AP Exam Tip: Keep the validity check separate from the fit comparison. A full-credit explanation identifies the outcomes, verifies each model’s probabilities, compares observed with expected counts or relative frequencies, and states which valid model is closer to the data. Use cautious language such as “fits these data better,” not “is proven true.”

Common Mistakes

  • Comparing an invalid proposal as if it were a model. Check the probability range and total first. A list that sums to more than 1 cannot be a probability distribution for the stated outcomes.
  • Using different outcome categories. Align the categories before comparing. A model for four separate bins cannot be compared directly with data that combine some bins unless the model probabilities are combined in the same way.
  • Comparing probabilities with counts without scaling. Compare model probabilities with relative frequencies, or multiply each model probability by \(n\) and compare expected counts with observed counts.
  • Expecting exact equality. Expected counts are long-run averages under the model. A valid model can produce observed counts that differ from those expectations.
  • Relying on one category or one total alone. Inspect the discrepancies across all categories. A total can summarize the comparison, but it should not hide a consistent mismatch in a particular outcome.
  • Claiming a model is true because it fits better. State that it fits the data considered better. A finite data set cannot guarantee the model’s assumptions or its future performance.

Key Takeaway

Comparing competing probability models requires two separate judgments. First, determine whether each proposal is a valid distribution for the same complete set of outcomes. Then compare the predictions of the valid models with observed relative frequencies or expected counts, category by category. A closer match is evidence of better descriptive fit to those data, not proof of truth.

Key takeaway: Check validity before fit. Compare observed outcomes with each valid model on the same scale and in the same categories, then describe the result cautiously and in context.

Check Your Understanding

Use the ideas in this tutorial to evaluate and compare probability models.

  1. A model assigns probabilities \(0.25, 0.25, 0.30,\) and \(0.20\) to four exhaustive outcomes. Is it valid? Explain both checks you use.
  2. In 80 trials, a model assigns probability \(0.15\) to outcome C. What expected count does the model predict for C?
  3. For 100 trials, a category has observed count 28. Model A predicts probability \(0.25\), and Model B predicts probability \(0.31\). Find the expected count under each model and identify which is closer for this category.
  4. Why should a model comparison use matching outcome categories and data collected under relevant conditions?
  5. If one valid model is closer to the observed counts than another, what is an appropriate conclusion—and what claim would be too strong?