What a Conditional Distribution by Column Describes
In Conditional Distributions by Row, you used row totals to describe the column categories within a specified row group. Now switch the direction: a conditional distribution by column describes the row categories within a specified column group. The group being conditioned on is named by the column.
For example, if a table classifies people by travel method in rows and distance from work in columns, asking “What percentage of people who live near work walk?” calls for a column percentage. “People who live near work” identifies the group, so the denominator is the total in the Near column. The answer describes the travel methods of people in that column group.
The denominator follows the group named in the question. “Of people who live near work” points to the Near column; “of people who live far from work” points to the Far column. A column percentage is not a share of the entire sample, and it is not calculated using a row total.
Calculate Column Percentages
Consider a fictional survey of 240 residents. Each resident is classified by primary travel method and distance from work. Travel method appears in the rows; distance from work appears in the columns.
| Primary travel method | Near | Middle distance | Far | Row total |
|---|---|---|---|---|
| Bus | 16 | 30 | 24 | 70 |
| Bicycle | 32 | 20 | 6 | 58 |
| Walk | 32 | 50 | 30 | 112 |
| Column total | 80 | 100 | 60 | 240 |
To find the percentage of residents in the Near group who travel by bus, divide the Bus-and-Near cell count by the Near column total:
In this fictional survey, 20% of the residents who live near work primarily travel by bus. The denominator is 80 because that is the total number of residents in the Near column. The question is about residents who live near work, not all surveyed residents and not everyone who takes the bus.
A complete conditional distribution by column includes the percentages for every row category within a column. For the Near column, 32 of 80 residents bicycle and 32 of 80 walk. The percentages are 40% and 40%. Together with the 20% who take the bus, these account for all residents in that column.
Read and Compare Complete Column Distributions
Calculate each cell percentage using the total at the bottom of its own column. The completed distributions for the travel survey are:
| Primary travel method | Near | Middle distance | Far |
|---|---|---|---|
| Bus | 20% | 30% | 40% |
| Bicycle | 40% | 20% | 10% |
| Walk | 40% | 50% | 50% |
| Total | 100% | 100% | 100% |
For example, the Middle-distance bus percentage is \(30/100\times100\%=30\%\). The Far-column bicycle percentage is \(6/60\times100\%=10\%\). Notice that the percentages add to 100% within each column: 20% + 40% + 40% for Near, 30% + 20% + 50% for Middle distance, and 40% + 10% + 50% for Far.
The column distributions make it possible to compare travel methods across distance groups. In this sample, 40% of residents in the Near group primarily bicycled, compared with 10% of residents in the Far group. The Near-group percentage was 30 percentage points higher. This comparison describes the residents surveyed; it does not establish that distance caused a particular travel method.
Worked Examples: Conditional Distributions by Column
Worked Example: What Percentage of Near-Work Residents Walk?
Use the travel survey table. Find the percentage of Near residents who walk and the percentage of Far residents who walk. Then compare the two percentages.
Identify the relevant cells and column totals: In the Near column, 32 residents walk out of 80 total residents. In the Far column, 30 residents walk out of 60 total residents.
Calculate each column percentage:
Interpret and compare: In this survey, 40% of residents who live near work primarily walked, compared with 50% of residents who live far from work. The Far-group percentage was \(50\%-40\%=10\) percentage points higher. Each denominator is the total of the distance group named in the comparison.
Using 240 as the denominator in both calculations would answer a different question: what percentage of the entire sample consisted of Near residents who walked, or Far residents who walked. The question asks about the percentage within each distance group, so the column totals are appropriate.
Worked Example: Complete Column Distributions for Alert Preferences
A fictional survey of 150 residents asks which alert format they prefer for community updates. The residents are classified by preferred format and whether they usually receive updates on a phone or a computer. Calculate the distribution of preferred format within each device group.
| Preferred alert format | Usually use a phone | Usually use a computer |
|---|---|---|
| Short text | 42 | 18 |
| Detailed message | 18 | 42 |
| Audio alert | 20 | 10 |
| Column total | 80 | 70 |
Calculate the phone column: Divide every cell in that column by 80.
The phone-group percentages total \(52.5\%+22.5\%+25\%=100\%\).
Calculate the computer column: Divide every cell in that column by 70.
The computer-group percentages total approximately \(25.7\%+60\%+14.3\%=100\%\). In this survey, 52.5% of residents who usually use a phone preferred short text, compared with about 25.7% of residents who usually use a computer. The phone-group percentage was about \(52.5\%-25.7\%=26.8\) percentage points higher. These distributions describe the residents surveyed, not all residents.
Worked Example: Four Steps for a Column-Percentage Comparison
A fictional community center surveys people about whether they attended a class during the past month. The table groups people by membership type in columns. Find and compare the percentage who attended in each membership group.
| Class attendance | Monthly pass | Drop-in |
|---|---|---|
| Attended | 54 | 24 |
| Did not attend | 18 | 24 |
| Column total | 72 | 48 |
We want to compare the percentage who attended a class among people with a monthly pass and among drop-in visitors.
For each membership group, divide the Attended cell count by that group’s column total and multiply by 100%. Each column total represents all surveyed people in that membership group.
Monthly pass: \(54/72\times100\%=75\%\). Drop-in: \(24/48\times100\%=50\%\). The difference is \(75\%-50\%=25\) percentage points.
In this survey, 75% of people with a monthly pass attended a class during the past month, compared with 50% of drop-in visitors. The monthly-pass percentage was 25 percentage points higher.
A check confirms that each column accounts for its full group. Among monthly-pass holders, \(54/72\times100\%=75\%\) attended and \(18/72\times100\%=25\%\) did not. Among drop-in visitors, \(24/48\times100\%=50\%\) attended and \(24/48\times100\%=50\%\) did not. Each column totals 100%. The table does not show that membership type caused the difference in attendance.
How the Conditioning Group Changes the Question
The same cell can be used for different questions, but the denominator changes when the group of interest changes. Consider the 54 people who both held a monthly pass and attended a class. There were 72 monthly-pass holders, 78 people who attended, and 120 people in the whole survey.
- \(54/72\times100\%=75\%\) answers: “What percentage of monthly-pass holders attended?” The conditioning group is the Monthly pass column.
- \(54/78\times100\%\approx69.2\%\) answers: “What percentage of people who attended had a monthly pass?” The conditioning group is the Attended row.
- \(54/120\times100\%=45\%\) answers: “What percentage of everyone surveyed both held a monthly pass and attended?” This is a joint percentage using the grand total.
These are not competing answers to the same question. They describe different groups. A column percentage conditions on a column category and describes how the row categories are distributed within that group. As in Conditional Distributions by Row, read the group phrase carefully before choosing a denominator. If the question instead names the row group, the row total is relevant.
Look for phrases such as “among phone users,” “of monthly-pass holders,” or “for people in the Far group.”
If the group is a column category, use that column. If it is a row category, a row percentage answers the within-group question.
For a column group, divide the relevant cell count by its column total. For a row group, divide by its row total.
Name the group in the denominator and the category in the cell. If comparing groups, state the percentage-point difference.
A completed column distribution adds to 100% down each column because those percentages account for every individual in that column group. The percentages across a row generally do not add to 100%: they use different column totals. Small deviations from 100% within a column can occur when percentages are rounded.
Common Mistakes and AP Exam Tips
- Using a row total for a column-group question. For “What percentage of phone users prefer short text?”, divide the Short text-and-phone count by the phone column total.
- Using the grand total as the denominator. The grand total gives a joint percentage, which describes a share of the entire sample. It does not describe a share within a column group.
- Assuming that column percentages add across a row. Each column percentage has a different column total. Check that percentages add to 100% down each column instead.
- Leaving out the conditioning group. “52.5% preferred short text” is unclear when the table contains phone and computer users. A complete interpretation says, “52.5% of surveyed residents who usually use a phone preferred short text.”
- Comparing counts when group sizes differ. The phone and computer groups may have different totals. Compare their within-column percentages when the question asks how common a preference is within each group.
- Calling a subtraction a percent difference. Subtract the two percentages and report the result in percentage points. For example, 75% versus 50% is a difference of 25 percentage points.
- Claiming that a group difference proves a cause. A two-way table describes how categories occur in the observed groups. A difference in conditional percentages alone does not show why it occurred.
For a strong response, show the cell count divided by the correct column total, convert to a percentage, and interpret the result by naming both categories. When comparing groups, report both percentages and their difference in percentage points.
Check Your Understanding
Use the travel-method and distance table in this tutorial. Show the denominator you choose and interpret each result in context.
- What percentage of residents in the Middle-distance group primarily travel by bus?
- What percentage of residents in the Far group primarily travel by bicycle?
- Write the complete conditional distribution of travel method for the Near group, and check its total.
- Compare the percentage who walk in the Near group with the percentage who walk in the Far group. State the difference in percentage points.
- For the 32 Near residents who bicycle, explain the different questions answered by dividing 32 by 80 and by 240.