Tutorials › AP Statistics › Why P(A | B) Is Not P(B | A)

Conditional probability · Tutorial 266 of 1000

Why P(A | B) Is Not P(B | A)

Learn to calculate both directions of a medical screening probability and explain why a positive result does not have the same probability as disease among people who test positive.

Beginner 9 min read

What You'll Learn

  • Identify the reference group in each direction of a conditional probability.
  • Distinguish the probability of a positive test among people with a condition from the probability of the condition among people with a positive test.
  • Organize screening information in a two-way table of counts.
  • Use prevalence, sensitivity, and false positive rate to calculate both conditional probabilities.
  • Explain how the proportion of people with the condition affects the probability that a positive result indicates the condition.

Two Questions That Sound Similar but Are Not

A medical screening test gives a result, such as positive or negative, and a person either has or does not have the condition being screened for. It is tempting to treat “the probability of testing positive if a person has the condition” as the same probability as “the probability that a person has the condition if the test is positive.” They are not the same question.

The notation makes the difference visible. Let \(D\) mean that a person has the condition and \(+\) mean that the person receives a positive test result. Then \(P(+\mid D)\) asks about test results among people who have the condition. In contrast, \(P(D\mid +)\) asks about condition status among people who test positive. The event after the bar sets the reference group, as in What Conditional Probability Means and Reading the Given Condition in a Sentence.

Definition: \(P(+\mid D)\) is the probability of a positive result among people who have the condition. \(P(D\mid +)\) is the probability of having the condition among people who test positive. The two probabilities reverse the target event and the condition, so their denominators—and generally their values—are different.

A test’s sensitivity is the proportion of people with the condition who test positive. In probability notation, sensitivity is \(P(+\mid D)\). A test’s false positive rate is the proportion of people without the condition who nevertheless test positive, \(P(+\mid D^c)\). Neither quantity directly answers \(P(D\mid +)\).

To find \(P(D\mid +)\), restrict attention to everyone who tested positive. That group includes people who have the condition and people who do not. The fraction of positive testers who have the condition is the count with both \(D\) and \(+\), divided by the total number of positive results. A two-way table makes that denominator easy to see.

$$ P(+\mid D)=\frac{\text{people with }D\text{ who test positive}}{\text{all people with }D} \qquad P(D\mid +)=\frac{\text{people with }D\text{ who test positive}}{\text{all people who test positive}} $$

The two calculations share the count of people who both have the condition and test positive. But one denominator counts everyone with the condition, while the other counts everyone with a positive result. That is why the probabilities cannot be swapped just because they use the same overlap.

Build a Table to See the Reversal

A practical approach is to imagine or calculate counts for a group of people. Begin with the proportion who have the condition, called the prevalence in this screening group. Then use the test’s sensitivity and false positive rate to find the positive counts in the condition and no-condition groups. The resulting table shows exactly which group belongs in each denominator.

The overall positive group combines positive results from people with the condition and positive results from people without it. In probability terms, these are disjoint groups, so their probabilities can be added. This uses the addition rule covered earlier in the course.

$$ P(D\mid +)= \frac{P(D\cap +)} {P(D\cap +)+P(D^c\cap +)} $$

This expression is another way to calculate a conditional probability: the numerator is the probability of both having the condition and testing positive; the denominator is the probability of testing positive. It does not say that a positive result causes the condition or that the test is necessarily useful in every setting. It answers a probability question for the population and rates specified.

Worked Example: A Rare Condition and a Positive Result

Worked Example: A Rare Condition and a Positive Result

Suppose a fictional screening program considers 10,000 people. In this hypothetical group, 1% have a particular condition. The test has a sensitivity of 90%, and 5% of people without the condition receive a positive result. Find \(P(+\mid D)\) and \(P(D\mid +)\), and explain why they differ.

State: Let \(D\) mean that a person has the condition and \(+\) mean a positive test. The first requested probability is \(P(+\mid D)\); the second is \(P(D\mid +)\).

Plan: Use the stated rates to build counts for the 10,000 people. One percent of the group has the condition, so 100 have it and 9,900 do not. The sensitivity determines the positive count among the 100 with the condition; the false positive rate determines the positive count among the 9,900 without it. Then use the proper reference group for each conditional probability.

Do: Among the 100 people with the condition, 90% test positive: \(0.90(100)=90\). Therefore, 10 test negative. Among the 9,900 people without the condition, 5% test positive: \(0.05(9{,}900)=495\). Therefore, 9,405 test negative. The complete table is:

PositiveNegativeTotal
Has condition9010100
Does not have condition4959,4059,900
Total5859,41510,000

For \(P(+\mid D)\), the reference group is the 100 people with the condition. For \(P(D\mid +)\), the reference group is the 585 people with positive results.

$$ P(+\mid D)=\frac{90}{100}=0.90 \qquad P(D\mid +)=\frac{90}{585}\approx 0.1538 $$

As a second check using probabilities, \(P(D\cap +)=90/10{,}000=0.009\), and \(P(+)=585/10{,}000=0.0585\). Thus \(P(D\mid +)=0.009/0.0585\approx0.1538\), matching the table calculation. The overall positive probability also checks from the two positive groups: \(0.009+0.0495=0.0585\).

Conclude: In this hypothetical screening group, 90% of people who have the condition test positive. But among people who test positive, about 15.4% have the condition. The condition is uncommon in this group, so the 495 positive results among people without it substantially outnumber the 90 positive results among people with it.

There is no contradiction between the answers. The first percentage uses people with the condition as its denominator; the second uses everyone with a positive result. Reversing the wording reverses the reference group.

Worked Example: The Same Test in a Different Population

Worked Example: The Same Test in a Different Population

Now imagine a fictional group of 1,000 people being screened in a setting where 10% have the condition. Suppose the test still has 90% sensitivity, but 10% of people without the condition test positive. Find both \(P(+\mid D)\) and \(P(D\mid +)\).

State: \(P(+\mid D)\) is the positive-result probability among people with the condition. \(P(D\mid +)\) is the condition probability among people with a positive result.

Plan: Ten percent of 1,000 is 100 people with the condition; the remaining 900 do not have it. Apply the sensitivity to the first group and the false positive rate to the second. The denominators will be 100 for the first conditional probability and the total positive count for the second.

Do: Of the 100 people with the condition, \(0.90(100)=90\) test positive and 10 test negative. Of the 900 without the condition, \(0.10(900)=90\) test positive and 810 test negative. Thus there are \(90+90=180\) positive results.

PositiveNegativeTotal
Has condition9010100
Does not have condition90810900
Total1808201,000
$$ P(+\mid D)=\frac{90}{100}=0.90 \qquad P(D\mid +)=\frac{90}{180}=0.50 $$

Check the reverse conditional with probabilities: \(P(D\cap +)=90/1{,}000=0.09\), while \(P(+)=180/1{,}000=0.18\). The ratio \(0.09/0.18=0.50\), the same result as \(90/180\).

Conclude: In this fictional group, 90% of people with the condition test positive, while 50% of positive testers have the condition. The sensitivity remained 90%, but the reverse conditional differs from the earlier example because the prevalence and false positive rate changed.

This comparison highlights why a test’s sensitivity alone cannot tell us the probability that a positive tester has the condition. To calculate that reverse probability, we must also account for how many people without the condition test positive and how common the condition is in the group.

Worked Example: A Negative Result Is Also Directional

Worked Example: A Negative Result Is Also Directional

Consider another fictional screening group of 1,000 people. Twenty percent have the condition. The test has 80% sensitivity, and 10% of people without the condition test positive. Find the probability of a negative result among people with the condition, \(P(-\mid D)\), and the probability of having the condition among people with a negative result, \(P(D\mid -)\).

State: The first question conditions on having the condition. The second conditions on receiving a negative result. They are different conditional probabilities, even though both refer to negative results or the condition.

Plan: There are 200 people with the condition and 800 without it. Since 80% of those with the condition test positive, the remaining 20% test negative. Of those without the condition, 10% test positive and 90% test negative. Use all 200 people with the condition as the denominator for \(P(-\mid D)\), and all negative testers as the denominator for \(P(D\mid -)\).

Do: Among those with the condition, \(0.20(200)=40\) test negative. Among those without it, \(0.90(800)=720\) test negative. The total number of negative results is \(40+720=760\).

PositiveNegativeTotal
Has condition16040200
Does not have condition80720800
Total2407601,000
$$ P(-\mid D)=\frac{40}{200}=0.20 \qquad P(D\mid -)=\frac{40}{760}\approx0.0526 $$

A probability-scale check gives \(P(D\cap -)=40/1{,}000=0.04\) and \(P(-)=760/1{,}000=0.76\). Their ratio is \(0.04/0.76\approx0.0526\), agreeing with the count calculation.

Conclude: In this hypothetical group, 20% of people with the condition receive a negative result. Among all people who receive a negative result, about 5.3% have the condition. These are not interchangeable statements: the first starts with people who have the condition, and the second starts with people who tested negative.

Common Mistakes and AP Exam Tips

  • Reversing the conditional bar. \(P(+\mid D)\) does not mean the same thing as \(P(D\mid +)\). Read the event after the bar as the group you are restricting to.
  • Using sensitivity as the probability of disease after a positive result. Sensitivity is \(P(+\mid D)\), not \(P(D\mid +)\). To find the latter, divide the number who both have the condition and test positive by the total number of positive testers.
  • Forgetting positive results among people without the condition. Those results belong in the denominator of \(P(D\mid +)\), even though those people do not belong in its numerator.
  • Reporting a probability without naming its reference group. A strong interpretation says, for example, “Among people who test positive, about 15.4% have the condition,” rather than only “The probability is 0.154.”
  • Assuming a test’s sensitivity determines the reverse conditional by itself. The reverse conditional also depends on the condition’s prevalence in the group and the positive-result rate among people without the condition.
  • Making a causal claim from the table. These calculations describe probabilities under the stated model; they do not show that the test causes the condition or that a positive result proves that a person has it.

For a clear solution, define the events, write each conditional probability in words, and identify its reference group before choosing the denominator. A table is especially useful because the overlap count can stay the same while the correct denominators change.

Key takeaway: \(P(+\mid D)\) asks what proportion of people with the condition test positive. \(P(D\mid +)\) asks what proportion of positive testers have the condition. The condition after the bar determines the reference group, so reversing the events generally changes the probability.

Check Your Understanding

Use the event after the bar to identify the reference group in each question.

  1. In a fictional group, 80 of 100 people with a condition test positive. What is \(P(+\mid D)\), and what group is its denominator?
  2. In that same group, 80 people with the condition and 120 people without it test positive. Find \(P(D\mid +)\). Explain why it is not the same as the answer to Question 1.
  3. A test has sensitivity 85%. In words, what conditional probability does this describe? Does it directly give the probability of having the condition after a positive result?
  4. Among 500 people without a condition, 25 test positive. What is the false positive rate? Which group is the reference group for this probability?
  5. Explain why knowing \(P(+\mid D)\) alone is not enough to calculate \(P(D\mid +)\).