Tutorials › AP Statistics › Evaluating Claims Based on Probability

Probability model interpretation · Tutorial 397 of 1000

Evaluating Claims Based on Probability

Use binomial probabilities to judge whether observed results would be unusual under a company’s claim, while keeping the conclusion appropriately cautious.

Intermediate 10 min read

What You'll Learn

  • Translate a company’s success-rate claim into a binomial model.
  • Choose an event that includes the observed result and relevant more extreme results.
  • Calculate a tail probability under the claimed model.
  • Explain how a small or large probability affects the plausibility of a claim.
  • Distinguish evidence against a claim from proof that it is false.
  • Match the direction of the probability calculation to the claim being evaluated.

What Does an Observed Result Say About a Claim?

A company may claim that a product works 90% of the time. After testing a small number of products, the observed success rate might be lower—or higher—than 90%. The difference alone does not tell us whether the claim is credible: random results vary from sample to sample. A useful question is, “If the company’s claim were true, how likely would it be to see a result this far from what the claim predicts?”

In Interpreting a Model Prediction in Context, you learned to describe a calculated probability as a chance under a stated model. Here, we use that idea to evaluate a claim. We temporarily treat the claim as true, calculate the probability of the observed result or a result at least as extreme in the relevant direction, and judge whether that probability is small enough to count as evidence against the claim.

Definition: A result is unusual under a probability claim when the model assigns a small probability to that result or to results at least as extreme. A small probability can be evidence against the claim, but it does not prove the claim false.

The phrase at least as extreme matters. If a company claims a high success rate and the observed result is unusually low, we generally count the observed result and lower success counts. If the question is about an exact success rate and results in either direction would challenge it, we may count extreme outcomes on both sides. State the event before calculating so the probability matches the question.

For a claimed success probability \(p\) over a fixed number \(n\) of independent trials, the number of successes can be modeled with a binomial random variable \(X\), provided the binomial conditions are reasonable. As discussed in Choosing Between Binomial and Other Models and Assumptions Behind a Probability Model, the model needs a fixed number of trials, two outcomes per trial, a constant success probability, and independent trials. Earlier tutorials explain how to check these features; below, each example states the relevant evidence.

Evaluation approach: Treat the claim as true for the calculation. Define the random variable, identify the observed-or-more-extreme event that fits the claim being evaluated, calculate its probability under the claimed model, and interpret how that probability affects the claim’s plausibility.

Worked Examples: Choosing and Interpreting a Tail

Worked Example: Testing a Claimed 90% Success Rate

A fictional company claims that 90% of its rechargeable lanterns pass a quality test. An inspector randomly selects 12 lanterns from a large production run and finds that 8 pass. Assess whether this result is unusual if the company’s claim is true.

State. Let \(X\) be the number of lanterns, among 12 selected, that pass the quality test. The company’s claim gives \(p=0.90\), so under the claim \(X\sim\operatorname{Binom}(12,0.90)\). The observed count is 8. A result at least as low as the observation is \(X\leq8\).

Plan. The 12 lanterns are a fixed number of trials, each lantern either passes or does not pass, and the claim assigns the same success probability, 0.90, to each lantern. The selection is random, and the production run is large enough that treating the outcomes as independent is reasonable. These conditions support the binomial model. Because the observed result is below the claimed rate, calculate the lower-tail probability \(P(X\leq8)\).

Do. It is convenient to count failures. If \(F\) is the number that fail, then \(F=12-X\), and \(X\leq8\) is equivalent to \(F\geq4\). Under the claim, \(F\sim\operatorname{Binom}(12,0.10)\). Use the complement of zero through three failures:

$$ P(X\leq8)=P(F\geq4)=1-P(F\leq3) $$

Expanding the cumulative probability gives:

$$ \begin{aligned} P(F\leq3) &=\binom{12}{0}(0.10)^0(0.90)^{12} +\binom{12}{1}(0.10)^1(0.90)^{11}\\ &\quad+\binom{12}{2}(0.10)^2(0.90)^{10} +\binom{12}{3}(0.10)^3(0.90)^9\\ &=0.2824295365+0.3765727153+0.2301277705+0.0852325076\\ &=0.9743625299 \end{aligned} $$

Therefore,

$$ P(X\leq8)=1-0.9743625299=0.0256374701\approx0.0256 $$

A calculator gives the same result as \(1-\operatorname{binomcdf}(12,0.10,3)\), approximately 0.0256. This calculation is under the claimed model; it is not the probability that the claim itself is true.

Conclude. If the company’s 90% success-rate claim is true and the model conditions hold, the chance that 12 randomly selected lanterns would include 8 or fewer that pass is about 0.0256, or 2.56%. That is a relatively small probability, so this result provides evidence against the claim. It does not prove the claim false: an unusual result can still occur by chance, and the model assumptions also matter.

Worked Example: Evaluating a Low Defect-Rate Claim

A fictional supplier claims that 5% of its glass bottles have a surface defect. A quality team randomly inspects 10 bottles from a much larger shipment and finds 3 with defects. Find the probability of 3 or more defects if the supplier’s claim is true, and interpret it.

State. Let \(D\) be the number of bottles with defects among the 10 inspected. If the claim is true, \(D\sim\operatorname{Binom}(10,0.05)\). The result of interest is \(D\geq3\), since three defects were observed and larger counts would be at least as unfavorable to the claim.

Plan. The inspection has a fixed 10 bottles, each bottle is classified as defective or not defective, and the claimed defect probability is 0.05 for each bottle. The bottles are randomly selected, and the shipment is large enough to treat their defect outcomes as independent. These conditions support the binomial calculation. Find the complement of zero, one, or two defects.

Do. Under the claim:

$$ P(D\geq3)=1-P(D\leq2) $$

Calculate the first three binomial probabilities:

$$ \begin{aligned} P(D\leq2) &=\binom{10}{0}(0.05)^0(0.95)^{10} +\binom{10}{1}(0.05)^1(0.95)^9\\ &\quad+\binom{10}{2}(0.05)^2(0.95)^8\\ &=0.5987369392+0.3151247049+0.0746347985\\ &=0.9884964426 \end{aligned} $$

Thus:

$$ P(D\geq3)=1-0.9884964426=0.0115035574\approx0.0115 $$

The calculator expression \(\operatorname{binomcdf}(10,0.05,2)\) gives about 0.9885 for the complement’s cumulative probability, so subtracting that value from 1 gives the same tail probability, about 0.0115.

Conclude. If the supplier’s 5% defect-rate claim is true, the chance of finding at least 3 defective bottles in a random inspection of 10 is about 0.0115, or 1.15%. This small probability is evidence against the claim. It does not establish the true defect rate, and the conclusion depends on the random selection and model assumptions being reasonable.

When Results on Either Side Matter

A one-direction claim points naturally to one tail. For example, if a company claims a success rate of at least 90%, unusually few successes are the results that challenge the claim. An unusually high count does not contradict a claim that success is at least 90%.

An exact claim can be different. If a company claims the success probability is exactly 50%, both unusually low and unusually high counts may cast doubt on that claim. The event used for evaluation should reflect the question’s meaning. For a symmetric binomial model with \(p=0.50\), counts equally far from the expected count are equally likely, which can make a two-direction calculation especially straightforward.

Worked Example: Checking an Exact 50% Claim in Either Direction

A fictional game-device company claims that a particular feature activates successfully on exactly half of attempts. A technician independently tests 10 attempts and observes 9 activations. Find the probability of an outcome at least as far from the model’s expected count as 9 is, in either direction.

State. Let \(X\) be the number of activations in 10 attempts. Under the exact claim, \(X\sim\operatorname{Binom}(10,0.50)\). The expected count is \(np=10(0.50)=5\). The observed count, 9, is 4 above 5. Counts at least as far from 5 are \(X\geq9\) or \(X\leq1\).

Plan. There are 10 fixed attempts, each has two possible outcomes, the claim assigns the same activation probability to every attempt, and the attempts are treated as independent. These facts support a binomial model. Because the claim is exactly 50% and either direction could challenge it, add the probabilities of the two equally distant tails.

Do. For \(X\geq9\), the possible counts are 9 and 10:

$$ P(X\geq9)=\binom{10}{9}(0.50)^9(0.50)^1+\binom{10}{10}(0.50)^{10} =\frac{10+1}{1024}=\frac{11}{1024}\approx0.0107 $$

By symmetry of the binomial model with \(p=0.50\), \(P(X\leq1)=P(X\geq9)\). Therefore, the probability of either tail is:

$$ P(X\leq1\text{ or }X\geq9) =2\left(\frac{11}{1024}\right) =\frac{22}{1024} \approx0.0215 $$

Conclude. If the company’s exact 50% claim is true, the chance of observing 9 or more activations or 1 or fewer is about 0.0215, or 2.15%. This is a relatively small probability and provides evidence against the exact claim. The two-tail calculation is appropriate here because the question treats unusually high and unusually low counts as departures from 50%. It would not automatically be appropriate for a claim that only sets a minimum success rate.

How to Communicate the Strength of the Result

The probability calculated in these examples is a conditional statement: it describes what could happen if the claim and the model assumptions are true. It is not the probability that the company is telling the truth, the probability that the observed result was caused by chance, or a guarantee about the next product tested.

A small tail probability means the observed result would not happen often under the claimed model. That makes the claim less plausible in light of the data, assuming the model represents the process well. A larger probability means the result is not especially surprising under the claim. It does not prove the claim correct; many possible claims may be compatible with a result that is not unusual.

Do not choose an event only because it produces a small probability. First decide what counts as a result that challenges the claim. For a high claimed success rate, the lower tail is generally relevant; for a low claimed defect rate, the upper tail is generally relevant. For an exact claim, decide whether departures in either direction matter and explain the choice.

  • Small probability: The result is unusual under the claim, so it provides evidence against the claim if the model conditions are credible.
  • Not-small probability: The result is plausible under the claim; there is not strong evidence against it from this result alone.
  • Either conclusion: Check the data collection and model assumptions before extending the conclusion to the real process.

Common Mistakes and AP Exam Communication

Evaluating a claim takes more than entering values into a calculator. A full-credit explanation connects the claim, the random variable, the event, and the probability in context.

  • Reporting only the probability of the exact count. If 8 successes were observed, \(P(X=8)\) leaves out even lower counts that also challenge a high success-rate claim. Explain why the event is \(X\leq8\), or otherwise justify the chosen event.
  • Using the wrong tail. A high success-rate claim is challenged by low success counts. A low defect-rate claim is challenged by high defect counts. Translate the context into an inequality before using a cumulative calculator command.
  • Treating the probability as the chance the claim is true. The calculated probability assumes the claim is true. It describes how often the chosen result would occur under that assumption, not the probability of the assumption itself.
  • Calling a small probability proof. A small probability is evidence against the claim, not certainty that it is false. Random variation and imperfect model assumptions remain relevant.
  • Ignoring what “extreme” means. For a claim that success is at least a stated rate, an unusually high result does not work against the claim. For an exact claim, both tails may matter. State the reasoning behind the event.
  • Leaving out context. A number such as 0.0115 is not a complete conclusion. Say what the 1.15% represents and under which claim and conditions it was calculated.
AP Exam Tip: Define \(X\), state the claimed model and the observed-or-more-extreme event, show the probability calculation, and finish with a contextual conclusion. Use “provides evidence against the claim” when the probability is small; do not say that the result proves the claim false.

Key Takeaway

A probability claim can be assessed by asking how often the observed result, or a result at least as extreme in the relevant direction, would occur if the claim were true. A small probability makes the result unusual under that model and provides evidence against the claim, but it does not prove the claim false. The event, the model conditions, and the interpretation all matter.

Key takeaway: Treat the claim as true for the calculation, choose an event that reflects how the observation challenges it, and interpret the resulting probability as evidence—not proof—about the claim.

Check Your Understanding

For each question, identify the relevant event and explain what the calculated probability would mean in context.

  1. A company claims that 80% of its filters pass a test. In a random test of 15 filters, 10 pass. What event would you calculate to assess this result against the claim? Explain why.
  2. A supplier claims that no more than 2% of its parts are defective. A random inspection finds 4 defective parts among 20. Which tail is relevant, and why?
  3. In your own words, explain why a probability of 0.01 under a company’s claim is not the probability that the claim is true.
  4. A company claims an exact success rate of 40%. Why might an evaluation consider unusually high and unusually low success counts?
  5. If the observed result is not unusual under the claimed model, does that prove the claim is correct? Explain.