Tutorials › AP Statistics › Reversing the Condition with Tree Diagrams

Conditional probability · Tutorial 274 of 1000

Reversing the Condition with Tree Diagrams

Use tree paths and conditional probability to find the chance that an outcome came from a particular first-stage category.

Beginner 9 min read

What You'll Learn

  • Identify the first-stage category and the observed outcome in a reversed conditional-probability question.
  • Find the complete-path probability for the category and outcome together.
  • Add all paths leading to the observed outcome to get the correct denominator.
  • Use a tree diagram to calculate \(P(\text{category}\mid\text{outcome})\).
  • Explain why reversing the order of the condition changes the reference group.

Reverse the Question, Not the Tree

A tree diagram usually begins with a category and then shows possible outcomes within that category. For example, a tree might first separate morning and evening trips, then show whether each trip is late. But a question may start with the outcome: among late trips, what proportion happened in the morning?

That asks for a conditional probability in the reverse direction. In Conditional Probability with Tree Diagrams, you learned to multiply branch probabilities to find the probability of a complete path. In Finding Total Probability from a Tree Diagram, you learned to add the paths that lead to the same outcome. Now combine those ideas: find the path for the category and outcome you want, then divide it by the total probability of the observed outcome.

Key idea: To find the probability of a first-stage category given a later outcome, divide the complete-path probability for that category and outcome by the total probability of the outcome across all first-stage categories.

Let \(A\) be a first-stage category and \(B\) the observed outcome. The complete path “\(A\), then \(B\)” has probability \(P(A)P(B\mid A)\). The conditional probability \(P(A\mid B)\) asks what fraction of all outcomes \(B\) came from category \(A\).

$$ P(A\mid B)=\frac{P(A\cap B)}{P(B)} =\frac{P(A)P(B\mid A)}{P(B)}. $$

If the first-stage categories \(A_1,A_2,\ldots,A_k\) are mutually exclusive and exhaustive, then every occurrence of \(B\) follows exactly one of those categories. Add the complete paths that end in \(B\) to get the denominator:

$$ P(A_i\mid B) = \frac{P(A_i)P(B\mid A_i)} {P(A_1)P(B\mid A_1)+P(A_2)P(B\mid A_2)+\cdots+P(A_k)P(B\mid A_k)}. $$

The denominator must be greater than 0, because it represents the probability of observing \(B\). This calculation does not change the tree’s branches. It changes which group you use as the reference group: for \(P(A\mid B)\), consider only the outcomes where \(B\) occurred.

Follow the Outcome Back Through the Tree

It can help to think of the paths ending in \(B\) as one group. The numerator is the part of that group that came from category \(A\); the denominator is the whole group. This is the same reference-group idea used in Conditional Probability from a Two-Way Table, but here the tree supplies the path probabilities.

1
Name the events.
Identify the first-stage category you want to know about and the outcome after the condition bar.
2
Find the target complete path.
Multiply the first-stage probability by the conditional probability of the outcome along that branch.
3
Find the total outcome probability.
Multiply along every path ending in the outcome, then add those path probabilities.
4
Divide and interpret.
Divide the target path probability by the total outcome probability, then describe the chance within the group defined by the condition.

The order of the events matters. \(P(B\mid A)\) describes the chance of outcome \(B\) among cases in category \(A\). In contrast, \(P(A\mid B)\) describes the chance of category \(A\) among cases with outcome \(B\). As explained in Why P(A | B) Is Not P(B | A), these probabilities generally differ.

Worked Example: Which Trips Were Late?

Worked Example: Which Trips Were Late?

In a fictional commute model, 60% of trips are made in the morning and 40% in the evening. The probability of being late is 0.10 for a morning trip and 0.25 for an evening trip. Given that a trip was late, find the probability it was a morning trip.

State: Let \(M\) mean that a trip was made in the morning, \(E\) mean that it was made in the evening, and \(L\) mean that the trip was late. The question asks for \(P(M\mid L)\), not \(P(L\mid M)\).

Plan: The morning and evening categories are mutually exclusive and exhaustive in this model, with probabilities \(0.60+0.40=1\). For each category, late and not late are complementary outcomes: the not-late probabilities are \(1-0.10=0.90\) in the morning and \(1-0.25=0.75\) in the evening. Find the morning-and-late path, then divide by the sum of both late paths. The total probability of late trips is positive, so it can be the denominator.

Do: Multiply along each path that ends in late:

$$ \begin{aligned} P(M\cap L)&=P(M)P(L\mid M)=0.60(0.10)=0.060,\\ P(E\cap L)&=P(E)P(L\mid E)=0.40(0.25)=0.100. \end{aligned} $$

Add those disjoint paths to find the total probability of a late trip:

$$ P(L)=0.060+0.100=0.160. $$

Now divide the morning-and-late path by all late trips:

$$ P(M\mid L)=\frac{P(M\cap L)}{P(L)} =\frac{0.060}{0.160}=0.375. $$

Conclude: In this model, given that a commute trip was late, the probability it was a morning trip is \(0.375\), or 37.5%.

Worked Example: A Positive Screening Result

Worked Example: A Positive Screening Result

Consider an invented screening model. Suppose 2% of people in a group have a certain disease. The probability of a positive test result among people with the disease is 0.95; among people without the disease, the probability of a positive result is 0.10. Given a positive result, find the probability that a person has the disease.

State: Let \(D\) mean that a person has the disease, and \(+\) mean that the test result is positive. The requested probability is \(P(D\mid +)\): the chance of disease within the group of people who tested positive.

Plan: The first-stage categories, disease and no disease, are mutually exclusive and exhaustive. Their probabilities are \(P(D)=0.02\) and \(P(D^c)=1-0.02=0.98\). For the no-disease category, the stated positive-result probability is \(P(+\mid D^c)=0.10\); the negative-result probability there is \(0.90\). For people with the disease, the negative-result probability is \(1-0.95=0.05\). Each pair of result branches sums to 1. Find the disease-and-positive path and divide by the total probability of a positive result. That total is positive.

Do: Calculate both paths that end in a positive result:

$$ \begin{aligned} P(D\cap +)&=P(D)P(+\mid D)=0.02(0.95)=0.019,\\ P(D^c\cap +)&=P(D^c)P(+\mid D^c)=0.98(0.10)=0.098. \end{aligned} $$

The total probability of a positive result is the sum of those paths:

$$ P(+)=0.019+0.098=0.117. $$

So the probability of disease given a positive result is:

$$ P(D\mid +)=\frac{P(D\cap +)}{P(+)} =\frac{0.019}{0.117}\approx 0.1624. $$

As a frequency check, imagine 10,000 people following this model. Then 200 have the disease, and \(200(0.95)=190\) of them test positive. The other 9,800 do not have the disease, and \(9{,}800(0.10)=980\) of them test positive. There are \(190+980=1{,}170\) positive results, of which 190 are among people with the disease. The proportion is \(190/1{,}170\approx0.1624\), matching the tree calculation.

Conclude: Under this invented model, given a positive test result, the probability that the person has the disease is about \(0.1624\), or 16.24%. This is a probability among people with positive results, not the probability of a positive result among people with the disease.

The Denominator Is the Whole Condition Group

In each example, the condition after the bar determined the denominator. For the commute question, the denominator included all late trips, whether morning or evening. For the screening question, it included all positive results, whether the person had the disease or not. This is why the denominator is found by adding every tree path that leads to the condition’s outcome.

A common shortcut is to use the first-stage probability as the denominator or to use just one branch probability. Neither describes the condition group. In the screening example, \(P(D)=0.02\) is the proportion of people with the disease in the full group, while \(P(+)=0.117\) is the proportion with a positive result. The question \(P(D\mid +)\) restricts attention to positive results, so it uses \(0.117\) as its denominator.

Worked Example: Which Supplier Was Involved?

Worked Example: Which Supplier Was Involved?

A fictional device company receives 70% of a component from Supplier A and 30% from Supplier B. The probability that a component has a connection problem is 0.03 for Supplier A’s components and 0.08 for Supplier B’s. Given that a randomly selected component has a connection problem, find the probability it came from Supplier B.

State: Let \(A\) and \(B\) mean that the component came from Supplier A or Supplier B, and let \(C\) mean that it has a connection problem. The requested probability is \(P(B\mid C)\).

Plan: The supplier categories are mutually exclusive and exhaustive, with \(0.70+0.30=1\). The no-problem branch probabilities are \(1-0.03=0.97\) for Supplier A and \(1-0.08=0.92\) for Supplier B, so each supplier’s two outcome branches sum to 1. Find both paths that end in a connection problem. Divide Supplier B’s path probability by the sum of those paths; the sum is positive.

Do: Multiply along the two connection-problem paths:

$$ \begin{aligned} P(A\cap C)&=0.70(0.03)=0.021,\\ P(B\cap C)&=0.30(0.08)=0.024. \end{aligned} $$

Thus the probability of a connection problem is \(0.021+0.024=0.045\), and the requested conditional probability is:

$$ P(B\mid C)=\frac{P(B\cap C)}{P(C)} =\frac{0.024}{0.045}\approx0.5333. $$

Conclude: In this model, given that a selected component has a connection problem, the probability it came from Supplier B is about \(0.5333\), or 53.33%.

Common Mistakes and AP Exam Tips

  • Reversing the condition without recalculating. \(P(+\mid D)\) and \(P(D\mid +)\) answer different questions. State in words which group is the reference group before calculating.
  • Using the wrong denominator. For \(P(A\mid B)\), use \(P(B)\), the total probability of the condition \(B\). Add every path leading to \(B\), not just the path in the numerator.
  • Dividing the branch probabilities directly. The numerator is a complete-path probability such as \(P(A)P(B\mid A)\), not merely the conditional branch probability \(P(B\mid A)\).
  • Leaving out a possible first-stage category. If the categories are not exhaustive, the denominator misses some outcomes satisfying the condition. Check that the first-stage probabilities sum to 1 and that the categories do not overlap.
  • Giving a number without interpreting it. A full-credit conclusion identifies both the condition group and the event whose probability is being described. For example: “Among people with positive results, the probability of disease is about 16.24%.”

A reliable written solution names the requested conditional probability, shows the complete path in its numerator, adds all paths that meet the condition for its denominator, and gives an interpretation in context. The next tutorial examines medical-testing outcomes more closely, but the same reverse-the-tree calculation applies whenever a later outcome can follow more than one first-stage category.

Key takeaway: To find \(P(A\mid B)\) from a tree, divide the complete-path probability \(P(A\cap B)\) by the total probability \(P(B)\). Add every path ending in \(B\) for the denominator, then interpret the result among cases where \(B\) occurred.

Check Your Understanding

For each question, identify the reference group before calculating.

  1. A transit model has morning trips with probability 0.70 and evening trips with probability 0.30. The probabilities of a delay are 0.05 and 0.20, respectively. Given a delay, find the probability the trip was in the evening.
  2. In the screening example, why is \(P(D\mid +)\) not equal to \(P(+\mid D)\)?
  3. A component comes from Supplier X with probability 0.80 and Supplier Y with probability 0.20. Problem probabilities are 0.04 for X and 0.10 for Y. Find the probability that a component with a problem came from Y.
  4. What paths belong in the denominator when finding \(P(A\mid B)\) from a tree?
  5. In a model, \(P(A)=0.25\), \(P(B\mid A)=0.40\), and \(P(B\mid A^c)=0.10\). Find \(P(A\mid B)\).