Tutorials › AP Statistics › Categorical or Quantitative Data: First Decision

Choosing a mean-inference procedure · Tutorial 761 of 1000

Categorical or Quantitative Data: First Decision

Classify the variable the question is about before choosing a procedure: categories lead to proportions, while numerical measurements lead to means.

Intermediate 9 min read

What You'll Learn

  • Identify the response variable that the question asks about.
  • Distinguish numerical measurements from category labels, even when labels use numbers.
  • Recognize when a yes-or-no outcome calls for a proportion rather than a mean.
  • Choose between a mean-based and proportion-based direction before considering the number of groups.
  • Explain why changing the question about the same data can change the relevant statistic.

Make the First Decision About the Variable

Before choosing an inference procedure, ask what kind of outcome the research question is about. Is it a numerical measurement, such as time or distance, or is it a category, such as “replied” or “did not reply”? This first decision helps you tell whether the question concerns a population mean or a population proportion.

This distinction matters even when a study includes groups or records several kinds of information. A study might record a person’s treatment group, age, and recovery status. The treatment group and recovery status are categories; age is quantitative. The procedure depends on the variable named in the question—not simply on how many groups there are or how many columns appear in the data.

Definition: A quantitative variable records a numerical measurement or count for each observational unit, where arithmetic on the values is meaningful. A categorical variable places each unit into one of a set of groups or labels. A question about the mean of a quantitative variable concerns an average measurement; a question about the proportion in a category concerns a share of units.

The inference sequence in this unit focuses on means. As you saw in “Writing Complete Conclusions for Mean Inference Problems,” mean inference can involve one population, paired differences, or two groups. Those settings differ in their design, but all concern a quantitative response. Before sorting out which mean procedure fits, make sure the response really is quantitative.

A Short Decision Routine

Start by identifying the observational unit—the person, object, or other entity for which a record is made. Then name the response variable: the outcome the question asks you to describe or compare. Ask what one recorded value means for one unit.

1
Find the outcome in the question.
Underline the result being measured or counted. Do not classify a group label if the question is about a different response.
2
Ask what one value represents.
Does it give a measurement or count with meaningful numerical differences, or does it identify a category?
3
Match the summary to the variable.
A quantitative response can be summarized by a mean. A categorical response can be summarized by a count or proportion in a category.
4
Only then consider the design.
If the question concerns a mean, later decisions include whether the data come from one sample, paired observations, or independent groups. The design does not turn categories into measurements.

A useful test is to ask whether subtracting two recorded values would have a meaningful interpretation. A difference of 12 seconds in task times is meaningful. A difference of 2 between the labels “bus” and “train” is not. The values might be stored as numbers, but storage format alone does not determine the variable type.

Categories, Codes, and Yes-or-No Outcomes

Categorical variables can have labels, such as “morning,” “afternoon,” and “evening,” or numerical codes, such as 1, 2, and 3. If the numbers merely stand for labels, calculating an average of the codes does not answer a meaningful question. For example, changing the codes from 1, 2, and 3 to 10, 20, and 30 would change that average without changing anyone’s category.

A yes-or-no response is categorical, even if it is recorded as 1 for “yes” and 0 for “no.” If the question asks what fraction of units answered “yes,” the relevant summary is the proportion answering “yes.” In a binary data set, the arithmetic mean of the 0s and 1s equals that sample proportion, but the question is still about a category’s proportion. That fact does not make the underlying response a quantitative measurement or make a mean procedure the natural choice.

Key distinction: A number can be a measurement, a count, or just a code for a category. Decide what the recorded value means in context. Do not choose a mean procedure merely because the data file contains numbers.

Sometimes one study includes both types of response. For instance, researchers might record whether each participant completed a program and how many days completion took. “Did the participant complete the program?” is categorical; “how many days until completion?” is quantitative. A question about completion rates concerns proportions, while a question about mean time among a specified set of participants concerns a mean. Read the question closely and classify the particular response it asks about.

Worked Examples: Classify the Response First

Worked Example: Which School Lunch Was Chosen?

A school records the lunch chosen by each of 50 students on a particular day. The choices are “pasta,” “salad,” and “sandwich.” The question is: “What proportion of these students chose salad?” Suppose 17 of the 50 students chose salad.

Identify the response: The response is each student’s lunch choice. The values are category labels, not measurements. The question focuses on the “salad” category.

Choose the kind of summary: Because the response is categorical and the question asks for the share in one category, calculate a proportion:

$$ \hat{p}=\frac{\text{number who chose salad}}{\text{number of students}} =\frac{17}{50}=0.34 $$

The sample proportion is 0.34, or 34%. This is a proportion question, not a question about the mean lunch choice. Coding pasta as 1, salad as 2, and sandwich as 3 would not make the lunch choices quantitative: the codes are labels, and their numerical differences have no meaningful interpretation.

Worked Example: Time to Load a Practice Quiz

A student measures the loading time, in seconds, for five attempts to open a practice quiz: 4.2, 3.8, 4.5, 5.1, and 4.4 seconds. The question is: “What is the average loading time across these attempts?”

Identify the response: The response is loading time in seconds. The difference between 5.1 seconds and 4.1 seconds is one second, so arithmetic differences are meaningful. Loading time is quantitative.

Choose the kind of summary: The question asks for an average measurement, so the relevant sample statistic is the mean. Add the five times and divide by the number of attempts:

$$ \bar{x}=\frac{4.2+3.8+4.5+5.1+4.4}{5} =\frac{22.0}{5}=4.4\text{ seconds} $$

The sample mean loading time is 4.4 seconds. The observations are numerical measurements, not categories. If the question instead asked whether each attempt loaded in under five seconds, the response for that new question would be a yes-or-no category, and the question would concern a proportion.

Worked Example: Comparing Mean Cycle Times by Machine

A workshop measures how many minutes each item takes to pass through a production step. Items are processed on Machine A or Machine B. A manager asks whether the true mean cycle time differs between items processed on the two machines. In a small illustrative sample, the times are 12, 15, and 13 minutes for Machine A, and 16, 14, and 15 minutes for Machine B.

Identify the response and groups: The response is cycle time in minutes, a quantitative measurement. The machine label is categorical; it identifies the group. The question is about the mean of the quantitative response within each machine group, not the proportion of items processed on either machine.

Check what the sample summaries describe: For Machine A, the sum is \(12+15+13=40\), so the sample mean is \(40/3\approx13.33\) minutes. For Machine B, the sum is \(16+14+15=45\), so the sample mean is \(45/3=15\) minutes. The observed difference in sample means, A minus B, is approximately \(13.33-15=-1.67\) minutes.

These sample calculations do not by themselves establish a population difference or select every detail of an inference procedure. They establish the direction of the first decision: this is a question about comparing means because the response being measured is quantitative. The group variable tells us how the measurements are organized; it does not change the response type.

Keep the Response Separate from the Grouping Variable

In a comparison, students sometimes classify the grouping variable instead of the outcome. A question might compare the mean number of minutes spent exercising by grade level. Grade level is categorical, but exercise time is quantitative. Since the question asks about the mean response, the analysis is about means. Conversely, a question comparing the proportion of students who bring lunch across grade levels concerns a categorical response, even though grade level is also categorical.

A quick way to avoid this mix-up is to complete two phrases: “The groups are defined by ___” and “The response measured for each unit is ___.” Then classify the second phrase. For a mean comparison, the response must be quantitative. If the response is a category, a question about the fraction in a category points toward proportions instead.

Decision check:
  • Response: What outcome does the research question ask about?
  • Type: Is each recorded response a meaningful numerical measurement or count, or a category?
  • Target summary: Does the question ask about an average measurement or a share of units in a category?
  • Organization: Are there groups, repeated measurements, or pairs? Consider this after classifying the response.

This order prevents a common shortcut: “There are two groups, so use a two-sample t test.” Two groups alone do not tell you that. If the response is quantitative and the groups are independent, a comparison of means may be relevant. If the response is categorical, the question may instead compare proportions. If the same units contribute related observations, the design must also be considered. Identifying the response is the first decision, not the entire procedure choice.

Common Mistakes and AP Exam Tips

  • Choosing from the number of groups: Two groups do not automatically mean a two-sample mean procedure. First classify the response, then consider whether the design involves independent groups or pairs.
  • Treating category codes as measurements: A code such as 1 for “online” and 2 for “in person” is a label. A mean of these codes depends on an arbitrary coding choice. A full-credit explanation says the outcome is categorical and identifies the relevant category or categories.
  • Confusing the group with the response: A study may group units by school, treatment, or machine while measuring a quantitative response. Name both variables and state which one the question asks you to summarize.
  • Calling every numerical outcome quantitative without context: Check what the numbers represent. Measurements and counts can be quantitative; identifiers, ranks used only as labels, and category codes are not automatically quantitative.
  • Calling a 0-or-1 outcome a mean question: If 1 means “yes” and 0 means “no,” a question about the percentage of yes responses is a proportion question. The numerical coding does not change its categorical meaning.
  • Assuming the first decision completes the procedure choice: Classifying the response tells you whether the question is about means or proportions. The study design and inference conditions still matter before carrying out a procedure.

On an AP response, make the reasoning explicit rather than writing only “use a t test” or “use a proportion.” For example: “The response is the number of minutes per item, which is quantitative, and the question asks about the mean cycle time, so this is a mean-inference setting.” Or: “The response is whether each customer renewed, a categorical outcome, and the question asks for the renewal rate, so it is a proportion setting.” These statements show that the procedure choice follows from the variable and the question.

Key takeaway: The first decision is about the response variable named in the research question. A meaningful numerical measurement or count points to a mean question; a category, including a yes-or-no outcome recorded with number codes, points to a proportion question. After that, use the study design to narrow the procedure.

Check Your Understanding

For each situation, identify the response variable, classify it as categorical or quantitative, and say whether the question concerns a mean or a proportion.

  1. A transit planner records each rider’s main way of getting to a station and asks what fraction arrive by bicycle. What is the response type and relevant summary?
  2. A technician records the mass, in grams, of each package and asks for the average mass. Why is this a mean question?
  3. A school records whether each student submitted a permission form and asks whether submission rates differ by grade. What is the response variable, and why does having grades as groups not make this a mean question?
  4. A survey records satisfaction as “low,” “medium,” or “high,” coded 1, 2, and 3. Why is the mean of the codes not automatically a meaningful average?
  5. A researcher compares the mean number of minutes spent reading per day for students in two clubs. Identify the grouping variable and the quantitative response.