Tutorials › AP Statistics › Decision Flowchart for Inference Scenarios

Statistical practices and exam synthesis · Tutorial 1009 of 1020

Decision Flowchart for Inference Scenarios

Follow the data and the question through a decision flowchart, then justify the inference procedure that fits each scenario.

Intermediate 9 min read

What You'll Learn

  • Trace a scenario from response type to parameter and inferential goal.
  • Distinguish one-proportion tests and intervals from comparisons of two proportions.
  • Tell independent quantitative groups from matched-pair measurements.
  • Separate chi-square tests of independence from tests of homogeneity.
  • Justify a procedure by describing the data structure and the question it answers.

Turn Procedure Choice Into a Repeatable Route

In “Choosing the Correct Inference Procedure,” you learned to sort inference questions by response type, parameter, and data structure. Here, the goal is to make that sorting process quick and dependable across a set of different scenarios. A decision flowchart helps you follow the same route each time instead of guessing from a familiar word in the story.

Begin with what was measured, not with the setting. A survey about a health program might produce a categorical yes-or-no response, while a study in the same setting might measure a quantitative number of hours. Then identify whether the question asks for an estimate or tests a claim, and whether the data come from one sample, independent groups, matched pairs, or a table of categorical variables.

Key idea: A strong procedure choice connects three things: the response type, the arrangement of the observations, and the population parameter or relationship in the question. A procedure name without that connection is only a guess.

Follow the Decision Flowchart

Use the questions below in order. You usually do not need to calculate anything to select a procedure. If a scenario does not state enough about sampling or assignment to settle a condition, identify the likely procedure from the question and data structure; assess its conditions separately.

1
Is the response categorical or quantitative?
A categorical response puts each individual into a category, such as completed or not completed. A quantitative response is a numerical measurement, such as distance, time, or temperature.
2
If it is categorical, what is being asked?
One binary response in one population leads to a one-proportion \(z\) procedure. A comparison of binary responses in two independent groups leads to a two-proportion \(z\) procedure. Counts for two categorical variables in one sample point to a chi-square test of independence; comparing a categorical distribution across groups or populations points to a chi-square test of homogeneity.
3
If it is quantitative, how are the observations arranged?
One sample leads to a one-sample \(t\) procedure. Two independent groups lead to a two-sample \(t\) procedure. If each measurement is linked to another measurement, use the within-pair differences and a paired \(t\) procedure.
4
Is the goal estimation or a test?
An interval estimates a population parameter; a test assesses evidence about a claim or comparison. Once the structure and goal are clear, name the corresponding procedure and parameter.

One distinction deserves special attention. A single sample classified by two categorical variables asks whether those variables are associated, which is the setting for a chi-square test of independence. Separate samples or groups classified by one categorical variable ask whether its distribution differs across the groups, which is the setting for a chi-square test of homogeneity. Both use tables of counts, but the study question determines which name fits.

The ten worked examples below use this route. Each explanation gives the procedure and the reason for it; condition checking is a separate part of planning, covered in the next tutorial.

Ten Scenarios: Name the Procedure and Explain Why

Worked Example: Estimating One Population Proportion

Scenario 1. A random sample of apartment residents is asked whether they have a working smoke alarm. The question asks for an estimate of the proportion of all residents in the city who do.

Decision. The response is categorical with two outcomes, and there is one sample from one population. Because the goal is to estimate a population proportion rather than test a claim, use a one-proportion \(z\) interval for \(p\), the proportion of city residents with a working smoke alarm. A mean procedure would not fit: the response is not a quantitative measurement.

Worked Example: Testing a Claim About One Proportion

Scenario 2. A school district samples students to ask whether they bring a reusable water bottle to school. Administrators want to know whether there is evidence that fewer than 60% of all district students do so.

Decision. There is one sample, one binary categorical response, and a claim about one population proportion. Because the question asks for evidence about a claimed value rather than an estimate, use a one-proportion \(z\) test for \(p\), with the alternative that \(p\) is less than 0.60. The direction comes from “fewer than.”

Worked Example: Comparing Two Proportions

Scenario 3. In a fictional experiment, users are randomly assigned to receive either a standard or redesigned app notification. The response is whether each user opens the notification within a day. The researchers ask whether the opening proportions differ.

Decision. The response is binary and the two groups contain different users. The question compares two population proportions, \(p_1\) and \(p_2\), so use a two-proportion \(z\) test for \(p_1-p_2\). This is not a paired procedure: each user contributes an outcome to only one group. “Differ” calls for a two-sided comparison, rather than a specified direction.

Worked Example: Estimating a Population Mean

Scenario 4. A random sample of community garden plots is used to estimate the average mass, in kilograms, of tomatoes harvested per plot during one season.

Decision. Harvest mass is quantitative, and the question asks for the average in one population based on one sample. The parameter is the population mean \(\mu\), so use a one-sample \(t\) interval for \(\mu\). The fact that there may be many tomato plants does not make this a proportion question; the response is each plot’s numerical harvest mass.

Worked Example: Testing a Claim About a Mean

Scenario 5. A random sample of rechargeable batteries is tested for operating time. A manufacturer’s stated average is 12 hours, and a consumer group asks whether the population mean operating time is less than 12 hours.

Decision. Operating time is quantitative, there is one sample, and the question tests a claim about one population mean. Use a one-sample \(t\) test for \(\mu\), with an alternative below 12 hours. A one-proportion procedure is inappropriate because the response is a measured time, not a yes-or-no category.

Worked Example: Comparing Independent Group Means

Scenario 6. Two independent random samples of commuters, one from a coastal town and one from an inland town, report their weekly commuting time in hours. The question asks whether the population means differ.

Decision. The response is quantitative, and the commuters in one sample are not matched to commuters in the other. The parameter is the difference between two population means, so use a two-sample \(t\) procedure, specifically a two-sample \(t\) test for the question of whether the means differ. Two locations do not automatically mean two-proportion inference; the measured outcome is hours.

Worked Example: Recognizing Matched Pairs

Scenario 7. Each participant in a fictional workplace study completes a typing task before and after using an ergonomic keyboard for several weeks. The response is words typed per minute, and the question asks whether the population mean changes.

Decision. Typing speed is quantitative, but the two measurements are linked because the same participant supplies both. Define a difference for each participant, such as after minus before, and use a paired \(t\) procedure for the population mean difference \(\mu_d\). Treating the two sets of scores as independent would discard the matching that defines the data structure.

Worked Example: Testing for Association Between Two Categorical Variables

Scenario 8. One random sample of museum visitors is classified by visit type (first visit or repeat visit) and by whether the visitor used an audio guide. The question asks whether visit type and audio-guide use are associated in the visitor population.

Decision. A single sample is classified by two categorical variables, and the question asks about an association between them. Use a chi-square test of independence. The table of counts describes combinations of categories; this is not a comparison of numerical means or an analysis of paired quantitative measurements.

Worked Example: Comparing Categorical Distributions Across Groups

Scenario 9. Separate random samples are taken from three neighborhoods. Each resident is classified by their preferred method for receiving local alerts: text, email, or phone call. The question asks whether the distribution of preferences is the same across neighborhoods.

Decision. The response is one categorical variable with three categories, and the study compares its distribution across separate populations. Use a chi-square test of homogeneity. The question is about whether the category proportions are alike across groups, not whether two categorical variables are associated within one sample.

Worked Example: Estimating a Difference Between Two Proportions

Scenario 10. Two independent random samples of customers are asked whether they prefer curbside pickup. A store manager wants an estimate of the difference in the population proportions who prefer it at two store locations.

Decision. The response is binary, and the samples are independent. The goal is to estimate \(p_1-p_2\), not to test a claim that the difference is zero. Use a two-proportion \(z\) interval. This is in the same procedure family as the test in Scenario 3, but the stated goal—an estimate—determines that an interval is needed here.

Common Mistakes and What a Clear Answer Says

  • Stopping at “two groups.” Two independent groups with numerical measurements suggest a two-sample \(t\) procedure; two groups with a binary response suggest a two-proportion \(z\) procedure. Name the response and the parameter, not just the number of groups.
  • Overlooking matching. Before-and-after values from the same individuals are linked. Say that the analysis uses within-pair differences and identifies \(\mu_d\); do not call the samples independent.
  • Confusing independence and homogeneity. For independence, one sample is classified by two categorical variables and the question concerns their association. For homogeneity, separate groups or populations are compared on one categorical variable’s distribution.
  • Ignoring the goal. “Estimate” calls for an interval, while “is there evidence” or “does it differ” calls for a test. Both can concern the same parameter, but they answer different questions.
  • Giving only a procedure name. A complete selection explains why the method fits. For example: “Because the same participants provide quantitative before-and-after scores, I would analyze their differences with a paired \(t\) procedure for \(\mu_d\).”
  • Trying to check every condition while still identifying the method. Select the procedure from the response, parameter, design, and goal first. Then check its relevant conditions explicitly, as in “Choosing the Correct Inference Procedure.”

For a concise AP response, state the procedure and connect it to the structure in one sentence. Mention the parameter or relationship when it helps make the reason unmistakable. Do not claim that a method’s conditions are satisfied merely because the procedure family has been identified.

Key takeaway: Follow the question from response type to data structure to goal. Categorical responses lead to proportion or chi-square procedures depending on the groups and variables; quantitative responses lead to \(t\) procedures, with paired measurements analyzed through differences. Finish by naming the parameter or relationship and explaining why the procedure matches it.

Check Your Understanding

For each situation, name the procedure and give one reason it fits the response and study structure.

  1. A random sample of hikers is asked whether they saw wildlife on a trail. The goal is to estimate the proportion of all hikers on that trail who did.
  2. Two independent groups of volunteers try different recipes, and the response is the number of minutes needed to prepare a meal. The question compares average preparation times.
  3. The same devices are tested before and after a software update, and the response is battery life in hours. The question asks whether the mean changes.
  4. One random sample of residents is classified by household type and preferred emergency-alert channel. The question asks whether the variables are associated.
  5. Separate random samples from several regions are classified by their preferred public-transport option. The question asks whether the preference distribution differs by region.