Tutorials › AP Statistics › Categorical Variables and Two-Way Tables

Chi-square tests for categorical data · Tutorial 541 of 1000

Categorical Variables and Two-Way Tables

Build and check two-way tables by organizing counts for two categorical variables and interpreting their cells and totals.

Intermediate 9 min read

What You'll Learn

  • Identify the two categorical variables represented in a two-way table.
  • Interpret a cell as the count in one combination of categories.
  • Calculate and verify row totals, column totals, and the grand total.
  • Distinguish the table’s two variables from the categories within each variable.
  • Explain why counts alone may not be enough to compare groups of different sizes.

One Table, Two Categorical Variables

In “Exam-Style Free Response on Two-Proportion Tests,” you worked with two groups and compared the proportions with a shared outcome. A broader way to organize categorical data is to record counts for combinations of categories from two variables. This organization is called a two-way table. It lets you see how the counts are distributed across both variables at once.

For example, a student survey could record each respondent’s class year and usual study location. Class year is one categorical variable; study location is another. Each student contributes a count to exactly one combination, such as “second-year student and library.” The table can then show the count in each combination, the total for each class year, the total for each location, and the overall number of students.

Definition: A two-way table organizes counts for two categorical variables. The interior entries, called cell counts, give the number of observations in each combination of categories. The row totals and column totals are called marginal totals; their name comes from their placement at the margins of the table.

The two variables are the characteristics recorded for each observation. Their categories are the possible groupings for those characteristics. For the study-location survey, “class year” and “study location” are the variables; “first-year,” “second-year,” “library,” and “café” are examples of categories. Keeping this distinction clear helps you label and interpret the table correctly.

How to Read the Cells and Margins

A table’s row and column headings identify the categories. The number at the intersection of a row and a column is the cell count for that particular combination. A row total adds the counts across one row, while a column total adds the counts down one column. The grand total is the number of observations represented in the whole table.

Suppose each surveyed student gives one class year and one usual study location. If the table has rows for class year and columns for location, a cell count answers a precise question: “How many students in this class year reported this location?” The row total answers “How many surveyed students are in this class year, across all listed locations?” The column total answers “How many surveyed students reported this location, across all listed class years?”

The orientation can be reversed: class year could be shown in columns and study location in rows. That changes where the totals appear, but not what the data mean. What matters is that the headings are clear and the same organization is used consistently throughout the table.

$$ \text{Row total}=\text{sum of the cell counts across that row} $$
$$ \text{Column total}=\text{sum of the cell counts down that column} $$
$$ \text{Grand total}=\text{sum of all cell counts} =\text{sum of the row totals} =\text{sum of the column totals} $$

These equalities provide a useful arithmetic check. If the row totals and column totals do not each add to the same grand total, revisit the entries or additions. A mismatch can indicate a calculation error, an omitted category, or a data-recording problem.

Worked Example: Class Year and Study Location

Worked Example: Class Year and Study Location

A fictional school surveys 120 students about their class year and usual study location. Each student chooses one location. The recorded counts are organized below.

Class yearLibraryDormCaféRow total
First-year18121040
Second-year15131240
Third-year1012830
Fourth-year43310
Column total474033120

The two categorical variables are class year and usual study location. The cell containing 13 is at the intersection of the second-year row and dorm column. It means that 13 surveyed second-year students reported the dorm as their usual study location. It does not mean that 13 students were surveyed in total from that row or that location.

To verify the second-year row total, add its three location counts: \(15+13+12=40\). The first-year total is \(18+12+10=40\); the third-year total is \(10+12+8=30\); and the fourth-year total is \(4+3+3=10\). Thus the row totals add to \(40+40+30+10=120\).

The library column total is \(18+15+10+4=47\). The dorm column total is \(12+13+12+3=40\), and the café column total is \(10+12+8+3=33\). The column totals add to \(47+40+33=120\), matching both the row-total sum and the stated number surveyed.

The table supports factual descriptions of these surveyed students. For instance, 47 of the 120 students reported the library, and 40 surveyed students were second-years. The table alone does not establish why students chose a location, or whether the counts represent all students at the school.

Building a Table from Records

When data begin as individual records, a reliable first step is to identify the two variables and list their categories. Then count each observation once in the cell that matches both of its recorded categories. Make sure every observation can be placed in one and only one cell. If some responses are missing or allow more than one category, decide how they will be handled and label that decision clearly.

The categories should cover the observations being summarized. For example, if a survey includes students who study in a location other than the three most common choices, the table should include an “Other” category or explain how those responses were treated. Leaving responses out without saying so can make the grand total misleading.

1
Name the variables.
State what two categorical characteristics were recorded for each observation.
2
List the categories.
Use clear row and column headings, and include categories needed to account for the observations being summarized.
3
Count each combination.
Place each observation in the cell for its category on the row variable and its category on the column variable.
4
Add and check the margins.
Calculate every row total and column total, then verify that both sets of totals agree on the grand total.

Worked Example: Wildlife Rescue Intake Counts

Worked Example: Wildlife Rescue Intake Counts

A fictional wildlife rescue center summarizes 60 recent animal intakes by animal type and primary intake reason. Each animal is assigned one type and one primary reason.

Animal typeInjuryIllnessOrphanedRow total
Bird143825
Mammal971127
Reptile5218
Column total28122060

The row and column variables are animal type and primary intake reason. The cell count 11 means that 11 of the recorded animals were mammals whose primary intake reason was orphaning. The mammal row total is \(9+7+11=27\), so 27 of the animals were mammals, regardless of the listed reason.

Adding down the injury column gives \(14+9+5=28\). The illness total is \(3+7+2=12\), and the orphaned total is \(8+11+1=20\). Those column totals sum to \(28+12+20=60\). The row totals also sum to \(25+27+8=60\), so both checks agree with the stated grand total.

The table describes counts for the intake records included in this summary. It does not, by itself, tell us whether one animal type is more likely than another to arrive for a particular reason. That kind of comparison needs to account for how many animals of each type are represented.

Counts Are Not the Same as Proportions

A count reports how many observations are in a category or cell. A proportion describes a count relative to a stated total. These are different summaries, and a larger count does not automatically mean a larger proportion. This matters especially when the row totals or column totals differ.

In the class-year table, 40 first-year students and 40 second-year students were surveyed, so those row totals are equal. But the fourth-year row contains only 10 students. Comparing raw counts across the fourth-year and first-year rows without considering those different totals can give an incomplete picture. A count table is the starting point; when a question asks how common something is within a group, the relevant group total also matters.

Similarly, a column total combines observations from all row categories. It answers how many observations fall in that column across the whole table; it is not the number in every row or the total for a particular subgroup. Always identify the denominator or group when describing a proportion. Later work with categorical data will use table counts to make more formal comparisons, but first the rows, columns, and totals must be read correctly.

Worked Example: Payment Method and Store Type

Worked Example: Payment Method and Store Type

A fictional shopper survey records the payment method used and the type of store visited. The survey includes 180 shopping visits.

Store typeCashCardMobile paymentRow total
Small market12281050
Grocery store18522090
Campus store8221040
Column total3810240180

The two variables are store type and payment method. The 52 in the grocery-store row and card column counts visits that match both categories. It is not the total number of card payments in the survey; the card column total is \(28+52+22=102\).

The row totals check as follows: \(12+28+10=50\) small-market visits, \(18+52+20=90\) grocery-store visits, and \(8+22+10=40\) campus-store visits. Their sum is \(50+90+40=180\). The column totals are \(12+18+8=38\) cash visits, \(28+52+22=102\) card visits, and \(10+20+10=40\) mobile-payment visits. They also sum to 180.

There are more card visits than cash visits in this survey: 102 compared with 38. That is a correct count comparison for the whole set of surveyed visits. To compare payment patterns within a specific store type, however, the store-type row total is relevant. The counts should not be described as proportions unless they are divided by an appropriate total.

Common Mistakes and AP Exam Tips

A table is easy to misread when the variable names, categories, or totals are treated casually. Avoid these errors:

  • Calling categories the variables: “Class year” is a variable; “first-year” is one of its categories. Name the characteristic and the category separately when interpreting a cell.
  • Reading a cell as a total: A cell count belongs to one row-and-column combination. Read both headings before describing what it counts.
  • Mixing up row and column totals: A row total adds across a row; a column total adds down a column. Use the headings and the displayed orientation rather than relying on memory.
  • Forgetting the grand-total check: Row totals and column totals must each add to the same overall count. If they do not, check the arithmetic and whether observations or categories were omitted.
  • Calling a count a proportion: “There were 47 library responses” is a count. A statement about the fraction or percent of surveyed students requires comparison with a specified total.
  • Assuming the table explains a pattern: A table summarizes recorded counts. It does not, by itself, prove that one variable caused another or that the pattern applies to a wider population.
AP Exam Tip: For a complete table interpretation, name both categories in the cell, use “row total,” “column total,” or “grand total” precisely, and show the addition when a total is requested. If you compare counts across groups of different sizes, point out that the counts alone do not compare within-group proportions.

Key Takeaway

A two-way table is a compact way to organize observations classified by two categorical variables. Each interior cell records one combination of categories; the margins summarize one variable at a time; and the grand total counts all observations represented. Accurate labels and matching totals make the table interpretable and provide a quick check on the data organization.

Key takeaway: Read a cell using both its row and column categories, calculate row totals across rows and column totals down columns, and confirm that both sets of margins add to the same grand total. Keep counts distinct from proportions and interpretations about a broader population.

Check Your Understanding

Use the table structure and terminology from this tutorial to answer each question.

  1. A table has “grade level” as its row variable and “preferred lunch” as its column variable. What does a cell count represent?
  2. In a two-way table, the counts in one row are 7, 11, and 6. Find that row total and explain what it summarizes.
  3. A table’s column totals are 19, 24, and 17. Find the grand total and state one way to check it using row totals.
  4. Why should you be careful when comparing cell counts for two groups with different row totals?
  5. A cell at the intersection of “mammal” and “illness” contains 7. Describe that count in a complete sentence, naming both categories.