Tutorials › AP Statistics › Simpson’s Paradox in Categorical Data

Categorical tables and summaries · Tutorial 36 of 1000

Simpson's Paradox in Categorical Data

See how a third categorical variable can change the story told by an overall percentage, and learn to compare the overall and subgroup patterns carefully.

Beginner 9 min read

What You'll Learn

  • Recognize Simpson’s paradox in a two-way table and subgroup tables
  • Calculate and compare outcome percentages using the correct group totals
  • Explain how a lurking variable can be related to both group membership and outcome
  • Distinguish an overall association from the patterns within subgroups
  • Avoid treating a descriptive reversal as proof of cause and effect

When an Overall Comparison Reverses

In Comparing Conditional Percentages in Context, you learned to compare percentages within the groups named by a question. This tutorial adds a caution: an overall comparison can point in the opposite direction from comparisons made within meaningful subgroups. The reversal is called Simpson’s paradox.

The paradox is not a calculation error. It can happen when the groups being compared contain different mixes of a third variable, and that third variable is related to the outcome. Looking only at the overall percentages can hide those differences in group composition.

Definition: Simpson’s paradox occurs when a relationship or trend seen in combined data is reversed in the corresponding subgroup comparisons after the data are separated into groups defined by a third variable. A variable that helps account for the pattern is called a lurking variable when it was not included in the original comparison.

For a clear comparison, keep the outcome category and denominator consistent. If the outcome is recovery, for example, compare the percentage who recovered out of everyone in each treatment group, first overall and then within each subgroup. A subgroup percentage uses that subgroup’s own total, as in the earlier tutorials on conditional distributions.

Worked Example: Recovery Rates and Patient Severity

Worked Example: Recovery Rates and Patient Severity

Suppose a fictional clinic records whether patients recovered after receiving Treatment A or Treatment B. It also records whether each patient’s condition was mild or severe at the start. The table shows invented counts; each patient is counted once.

Condition groupTreatmentRecoveredDid not recoverTotal
MildA721890
MildB18220
SevereA41620
SevereB92130

Start by comparing recovery rates among patients with mild conditions. Use each treatment’s total in the mild group as the denominator:

$$ \text{A, mild: }\frac{72}{90}\times100\%=80\%, \qquad \text{B, mild: }\frac{18}{20}\times100\%=90\%. $$

Among patients with mild conditions, Treatment B’s recovery rate is 10 percentage points higher than Treatment A’s. Now compare the rates among patients with severe conditions:

$$ \text{A, severe: }\frac{4}{20}\times100\%=20\%, \qquad \text{B, severe: }\frac{9}{30}\times100\%=30\%. $$

Among patients with severe conditions, Treatment B’s recovery rate is also 10 percentage points higher. Thus, within each severity group, Treatment B has the higher observed recovery rate.

Next, combine the severity groups for each treatment. Treatment A has \(72+4=76\) recoveries among \(90+20=110\) patients. Treatment B has \(18+9=27\) recoveries among \(20+30=50\) patients:

$$ \text{A overall: }\frac{76}{110}\times100\%\approx69.1\%, \qquad \text{B overall: }\frac{27}{50}\times100\%=54\%. $$

Overall, Treatment A’s observed recovery rate is about \(69.1\%-54\%=15.1\) percentage points higher. That reverses the comparison within each severity group, where Treatment B’s rate was higher.

The group composition helps explain why these summaries differ. Among Treatment A’s patients, \(20/110\approx18.2\%\) had severe conditions. Among Treatment B’s patients, \(30/50=60\%\) had severe conditions. Since recovery rates are lower for patients with severe conditions in both treatment groups, the overall percentage is affected by the fact that Treatment B was used with a much larger share of severe cases in this table.

Conclusion: In these fictional data, Treatment B has the higher observed recovery percentage within both severity groups, while Treatment A has the higher overall percentage. Patient severity is a lurking variable for the overall comparison because it is related to both the treatment-group mix and recovery. These descriptive data alone do not establish that either treatment caused better outcomes.

Why the Mix of Groups Matters

An overall percentage combines both the outcomes in a group and the proportions of that group in its subcategories. If one treatment group contains mostly people in a subgroup with higher outcome rates, its overall rate can be high even if its rate is lower within every subgroup. Conversely, a group that contains more people in a subgroup with lower outcome rates can have a low overall percentage, even if its within-subgroup percentages are higher.

This is why a third variable can be important even when the first table appears to compare only two variables. In the clinic example, the two-way comparison is treatment and recovery. Adding severity reveals separate conditional comparisons and shows that the treatment groups do not have the same mix of mild and severe cases.

The word “lurking” does not mean that the third variable is hidden in every sense. It means that it was left out of the original comparison and may help account for the observed pattern. A variable is especially worth checking when it is plausibly related both to which group an individual is in and to the outcome being summarized.

Key comparison: An overall percentage answers a question about all individuals in a group, with that group’s actual subgroup mix. A within-subgroup percentage answers a question about individuals in one specified subgroup. These are different summaries, so they can point in different directions without either calculation being wrong.

Worked Example: Inspection Results by Job Difficulty

Worked Example: Inspection Results by Job Difficulty

A fictional building service compares the percentage of jobs that pass inspection for two installation methods. Jobs are also classified as easy or difficult. Find the overall and within-category pass rates and describe the reversal.

Job categoryMethodPassedDid not passTotal
EasyA72880
EasyB19120
DifficultA83240
DifficultB71320

Among easy jobs, the pass rates are:

$$ \text{A: }\frac{72}{80}\times100\%=90\%, \qquad \text{B: }\frac{19}{20}\times100\%=95\%. $$

Among difficult jobs, the rates are:

$$ \text{A: }\frac{8}{40}\times100\%=20\%, \qquad \text{B: }\frac{7}{20}\times100\%=35\%. $$

Method B has the higher pass rate in both categories: 5 percentage points higher for easy jobs and 15 percentage points higher for difficult jobs. Now add the counts for each method. Method A has \(72+8=80\) jobs passing out of \(80+40=120\). Method B has \(19+7=26\) passing out of \(20+20=40\):

$$ \text{A overall: }\frac{80}{120}\times100\%\approx66.7\%, \qquad \text{B overall: }\frac{26}{40}\times100\%=65\%. $$

The overall comparison favors Method A by about \(66.7\%-65\%=1.7\) percentage points, despite Method B having the higher rate within both job categories. The proportions of difficult jobs also differ: \(40/120\approx33.3\%\) of Method A’s jobs are difficult, compared with \(20/40=50\%\) of Method B’s jobs.

Conclusion: The overall table favors Method A, but the comparisons within both easy and difficult jobs favor Method B. Job difficulty is a plausible lurking variable because the methods were used on different mixes of jobs, and difficulty is related to the chance of passing inspection. The table describes these jobs; it does not by itself prove that one method caused a higher or lower pass rate.

Worked Example: Finding the Reversal in Training Data

Worked Example: Finding the Reversal in Training Data

A fictional workplace compares the percentage of learners who complete a training course using Platform X or Platform Y. Learners are also classified by prior experience. Use the table to identify the overall pattern, check both experience groups, and explain a possible lurking variable.

Prior experiencePlatformCompletedDid not completeTotal
ExperiencedX44650
ExperiencedY18220
New to the topicX62430
New to the topicY92130

Among experienced learners, the completion percentages are:

$$ \text{X: }\frac{44}{50}\times100\%=88\%, \qquad \text{Y: }\frac{18}{20}\times100\%=90\%. $$

Among learners new to the topic, they are:

$$ \text{X: }\frac{6}{30}\times100\%=20\%, \qquad \text{Y: }\frac{9}{30}\times100\%=30\%. $$

Platform Y has the higher completion rate in each experience group. Overall, Platform X has \(44+6=50\) completions out of \(50+30=80\) learners, while Platform Y has \(18+9=27\) completions out of \(20+30=50\) learners:

$$ \text{X overall: }\frac{50}{80}\times100\%=62.5\%, \qquad \text{Y overall: }\frac{27}{50}\times100\%=54\%. $$

The overall comparison favors X by \(62.5\%-54\%=8.5\) percentage points, reversing the within-group pattern. The experience mix differs: \(30/80=37.5\%\) of Platform X’s learners are new to the topic, compared with \(30/50=60\%\) of Platform Y’s learners. Since the completion percentages are lower for new learners on both platforms, prior experience may help explain the reversal.

Conclusion: Platform Y has higher observed completion percentages among experienced learners and among new learners, while Platform X has the higher overall percentage. Prior experience is a possible lurking variable in the overall comparison. These data do not establish that changing platforms would cause a learner’s chance of completion to increase or decrease.

Common Mistakes and AP Exam Tips

  • Using the grand total for a conditional percentage. A rate “among difficult jobs” uses the total number of difficult jobs for that method, not the total for all jobs. State the denominator group before dividing.
  • Reporting only the overall comparison. In a possible reversal, the subgroup percentages are essential. A complete description compares the outcome within each subgroup and then compares the combined groups.
  • Calling the paradox a contradiction or arithmetic mistake. The overall and subgroup percentages answer different questions and can both be correct. Explain how the subgroup mix differs rather than choosing one calculation as the “right” one.
  • Assuming the lurking variable proves a cause. A table may show that a third variable is related to group membership and outcome, but that alone does not show that it caused the reversal or that changing groups would change the outcome.
  • Switching the outcome category midway. If the comparison starts with the percentage who recovered, passed, or completed, use that same category in the overall and subgroup calculations.
  • Ignoring counts and group sizes. Give the relevant numerator and denominator or enough information to identify them. This makes it possible to check that percentages are based on the proper groups, as emphasized in Interpreting Percentages From Small Counts.

For full-credit communication, name the overall direction, describe what happens within each subgroup, and identify the third variable that may help explain the reversal. Use cautious wording such as “is associated with” or “may help explain.” Do not claim that a descriptive table proves a causal effect.

Key takeaway: Simpson’s paradox is a reversal between an overall comparison and comparisons within subgroups. Check the conditional percentages and subgroup sizes, then explain how a third variable related to both group composition and outcome may help account for the difference.

Check Your Understanding

For each question, focus on which group supplies the denominator and whether the overall pattern matches the subgroup patterns.

  1. A program has 12 successes among 15 experienced participants and 6 among 10 new participants. What are the success percentages in each group? State the denominator used for each.
  2. In a fictional table, Group A has a 70% success rate and Group B has a 60% success rate overall. After splitting by experience, Group B has a higher success rate in both groups. What feature of the data could allow both statements to be true?
  3. Why is it important to compare the same outcome category—such as “completed”—both overall and within each subgroup?
  4. A two-way table shows a reversal after splitting by job difficulty. Give one reason difficulty might be a lurking variable, and one reason the table alone does not establish causation.
  5. Write a short description of Simpson’s paradox that distinguishes an overall percentage from a within-subgroup percentage.