Tutorials › AP Statistics › Conditional Distributions by Column

Categorical tables and summaries · Tutorial 28 of 1000

Conditional Distributions by Column

Practice calculating column percentages and interpreting them as distributions within the groups named by the table’s columns.

Beginner 9 min read

What You'll Learn

  • Identify the column group that a conditional distribution describes.
  • Divide each cell count by its column total to calculate a column percentage.
  • Calculate a complete conditional distribution for each column group.
  • Check that percentages within a column add to 100%, allowing for rounding.
  • Compare percentages across column groups using percentage points.
  • Explain how changing the conditioning group changes the question answered.

What a Conditional Distribution by Column Describes

In Conditional Distributions by Row, you used row totals to describe the column categories within a specified row group. Now switch the direction: a conditional distribution by column describes the row categories within a specified column group. The group being conditioned on is named by the column.

For example, if a table classifies people by travel method in rows and distance from work in columns, asking “What percentage of people who live near work walk?” calls for a column percentage. “People who live near work” identifies the group, so the denominator is the total in the Near column. The answer describes the travel methods of people in that column group.

Definition: A column percentage is a cell count divided by its column total, expressed as a percentage. It describes the percentage of individuals in that column category who belong to the row category named by the cell.
Formula: Column percentage = cell count ÷ column total × 100%. Use the total for the specific column containing the cell.
$$ \text{Column percentage} = \frac{\text{count in the cell}}{\text{total for that column}} \times 100\% $$

The denominator follows the group named in the question. “Of people who live near work” points to the Near column; “of people who live far from work” points to the Far column. A column percentage is not a share of the entire sample, and it is not calculated using a row total.

Calculate Column Percentages

Consider a fictional survey of 240 residents. Each resident is classified by primary travel method and distance from work. Travel method appears in the rows; distance from work appears in the columns.

Primary travel methodNearMiddle distanceFarRow total
Bus16302470
Bicycle3220658
Walk325030112
Column total8010060240

To find the percentage of residents in the Near group who travel by bus, divide the Bus-and-Near cell count by the Near column total:

$$ \frac{16}{80}\times100\%=20\% $$

In this fictional survey, 20% of the residents who live near work primarily travel by bus. The denominator is 80 because that is the total number of residents in the Near column. The question is about residents who live near work, not all surveyed residents and not everyone who takes the bus.

A complete conditional distribution by column includes the percentages for every row category within a column. For the Near column, 32 of 80 residents bicycle and 32 of 80 walk. The percentages are 40% and 40%. Together with the 20% who take the bus, these account for all residents in that column.

Read and Compare Complete Column Distributions

Calculate each cell percentage using the total at the bottom of its own column. The completed distributions for the travel survey are:

Primary travel methodNearMiddle distanceFar
Bus20%30%40%
Bicycle40%20%10%
Walk40%50%50%
Total100%100%100%

For example, the Middle-distance bus percentage is \(30/100\times100\%=30\%\). The Far-column bicycle percentage is \(6/60\times100\%=10\%\). Notice that the percentages add to 100% within each column: 20% + 40% + 40% for Near, 30% + 20% + 50% for Middle distance, and 40% + 10% + 50% for Far.

The column distributions make it possible to compare travel methods across distance groups. In this sample, 40% of residents in the Near group primarily bicycled, compared with 10% of residents in the Far group. The Near-group percentage was 30 percentage points higher. This comparison describes the residents surveyed; it does not establish that distance caused a particular travel method.

Key check: In a complete conditional distribution by column, percentages down each individual column add to 100%, apart from small differences caused by rounding. Each column is calculated separately using its own total.

Worked Examples: Conditional Distributions by Column

Worked Example: What Percentage of Near-Work Residents Walk?

Use the travel survey table. Find the percentage of Near residents who walk and the percentage of Far residents who walk. Then compare the two percentages.

Identify the relevant cells and column totals: In the Near column, 32 residents walk out of 80 total residents. In the Far column, 30 residents walk out of 60 total residents.

Calculate each column percentage:

$$ \text{Near: } \frac{32}{80}\times100\%=40\% $$
$$ \text{Far: } \frac{30}{60}\times100\%=50\% $$

Interpret and compare: In this survey, 40% of residents who live near work primarily walked, compared with 50% of residents who live far from work. The Far-group percentage was \(50\%-40\%=10\) percentage points higher. Each denominator is the total of the distance group named in the comparison.

Using 240 as the denominator in both calculations would answer a different question: what percentage of the entire sample consisted of Near residents who walked, or Far residents who walked. The question asks about the percentage within each distance group, so the column totals are appropriate.

Worked Example: Complete Column Distributions for Alert Preferences

A fictional survey of 150 residents asks which alert format they prefer for community updates. The residents are classified by preferred format and whether they usually receive updates on a phone or a computer. Calculate the distribution of preferred format within each device group.

Preferred alert formatUsually use a phoneUsually use a computer
Short text4218
Detailed message1842
Audio alert2010
Column total8070

Calculate the phone column: Divide every cell in that column by 80.

$$ \frac{42}{80}\times100\%=52.5\%,\qquad \frac{18}{80}\times100\%=22.5\%,\qquad \frac{20}{80}\times100\%=25\% $$

The phone-group percentages total \(52.5\%+22.5\%+25\%=100\%\).

Calculate the computer column: Divide every cell in that column by 70.

$$ \frac{18}{70}\times100\%\approx25.7\%,\qquad \frac{42}{70}\times100\%=60\%,\qquad \frac{10}{70}\times100\%\approx14.3\% $$

The computer-group percentages total approximately \(25.7\%+60\%+14.3\%=100\%\). In this survey, 52.5% of residents who usually use a phone preferred short text, compared with about 25.7% of residents who usually use a computer. The phone-group percentage was about \(52.5\%-25.7\%=26.8\) percentage points higher. These distributions describe the residents surveyed, not all residents.

Worked Example: Four Steps for a Column-Percentage Comparison

A fictional community center surveys people about whether they attended a class during the past month. The table groups people by membership type in columns. Find and compare the percentage who attended in each membership group.

Class attendanceMonthly passDrop-in
Attended5424
Did not attend1824
Column total7248
1
State.
We want to compare the percentage who attended a class among people with a monthly pass and among drop-in visitors.
2
Plan.
For each membership group, divide the Attended cell count by that group’s column total and multiply by 100%. Each column total represents all surveyed people in that membership group.
3
Do.
Monthly pass: \(54/72\times100\%=75\%\). Drop-in: \(24/48\times100\%=50\%\). The difference is \(75\%-50\%=25\) percentage points.
4
Conclude.
In this survey, 75% of people with a monthly pass attended a class during the past month, compared with 50% of drop-in visitors. The monthly-pass percentage was 25 percentage points higher.

A check confirms that each column accounts for its full group. Among monthly-pass holders, \(54/72\times100\%=75\%\) attended and \(18/72\times100\%=25\%\) did not. Among drop-in visitors, \(24/48\times100\%=50\%\) attended and \(24/48\times100\%=50\%\) did not. Each column totals 100%. The table does not show that membership type caused the difference in attendance.

How the Conditioning Group Changes the Question

The same cell can be used for different questions, but the denominator changes when the group of interest changes. Consider the 54 people who both held a monthly pass and attended a class. There were 72 monthly-pass holders, 78 people who attended, and 120 people in the whole survey.

  • \(54/72\times100\%=75\%\) answers: “What percentage of monthly-pass holders attended?” The conditioning group is the Monthly pass column.
  • \(54/78\times100\%\approx69.2\%\) answers: “What percentage of people who attended had a monthly pass?” The conditioning group is the Attended row.
  • \(54/120\times100\%=45\%\) answers: “What percentage of everyone surveyed both held a monthly pass and attended?” This is a joint percentage using the grand total.

These are not competing answers to the same question. They describe different groups. A column percentage conditions on a column category and describes how the row categories are distributed within that group. As in Conditional Distributions by Row, read the group phrase carefully before choosing a denominator. If the question instead names the row group, the row total is relevant.

1
Name the group in the question.
Look for phrases such as “among phone users,” “of monthly-pass holders,” or “for people in the Far group.”
2
Locate that group in the table.
If the group is a column category, use that column. If it is a row category, a row percentage answers the within-group question.
3
Choose the matching total.
For a column group, divide the relevant cell count by its column total. For a row group, divide by its row total.
4
Interpret in context.
Name the group in the denominator and the category in the cell. If comparing groups, state the percentage-point difference.

A completed column distribution adds to 100% down each column because those percentages account for every individual in that column group. The percentages across a row generally do not add to 100%: they use different column totals. Small deviations from 100% within a column can occur when percentages are rounded.

Common Mistakes and AP Exam Tips

  • Using a row total for a column-group question. For “What percentage of phone users prefer short text?”, divide the Short text-and-phone count by the phone column total.
  • Using the grand total as the denominator. The grand total gives a joint percentage, which describes a share of the entire sample. It does not describe a share within a column group.
  • Assuming that column percentages add across a row. Each column percentage has a different column total. Check that percentages add to 100% down each column instead.
  • Leaving out the conditioning group. “52.5% preferred short text” is unclear when the table contains phone and computer users. A complete interpretation says, “52.5% of surveyed residents who usually use a phone preferred short text.”
  • Comparing counts when group sizes differ. The phone and computer groups may have different totals. Compare their within-column percentages when the question asks how common a preference is within each group.
  • Calling a subtraction a percent difference. Subtract the two percentages and report the result in percentage points. For example, 75% versus 50% is a difference of 25 percentage points.
  • Claiming that a group difference proves a cause. A two-way table describes how categories occur in the observed groups. A difference in conditional percentages alone does not show why it occurred.

For a strong response, show the cell count divided by the correct column total, convert to a percentage, and interpret the result by naming both categories. When comparing groups, report both percentages and their difference in percentage points.

Key takeaway: A conditional distribution by column describes the row categories within a specified column group. Divide each cell by its own column total, interpret the percentage as “of that column group,” and compare groups using percentage points.

Check Your Understanding

Use the travel-method and distance table in this tutorial. Show the denominator you choose and interpret each result in context.

  1. What percentage of residents in the Middle-distance group primarily travel by bus?
  2. What percentage of residents in the Far group primarily travel by bicycle?
  3. Write the complete conditional distribution of travel method for the Near group, and check its total.
  4. Compare the percentage who walk in the Near group with the percentage who walk in the Far group. State the difference in percentage points.
  5. For the 32 Near residents who bicycle, explain the different questions answered by dividing 32 by 80 and by 240.