Tutorials › AP Statistics › Mapping Course Topics to Exam Questions

Statistical practices and exam synthesis · Tutorial 1003 of 1020

Mapping Course Topics to Exam Questions

Connect the first five AP Statistics units to exam tasks, and practice identifying which concepts and evidence a question calls for.

Intermediate 10 min read

What You'll Learn

  • Identify the correct names and content of Units 1–5, including Sampling Distributions as Unit 5.
  • Match one-variable and two-variable data prompts to useful graphs, summaries, and interpretations.
  • Recognize how sampling plans and experiments support different kinds of conclusions.
  • Connect probability models and sampling distributions to questions about chance variation.
  • Use a flexible map for four free-response questions and multiple-choice sets without treating either as a fixed unit-by-unit sequence.
  • Distinguish descriptive regression in Unit 2 from inference for regression slopes in Unit 9.

Map the Statistical Work, Not Just the Unit Number

In “Four Statistical Practices in Context” and “Formulating Questions and Planning Data Collection,” you connected a statistical question to the way data are collected, analyzed, and interpreted. On an exam, that chain may appear in one question or be spread across several prompts. A useful first move is to identify the statistical work the prompt asks you to do, then connect it to the relevant unit.

The first five AP Statistics units build a foundation for analyzing data and reasoning about chance. They do not correspond one-to-one with four free-response questions. A question can combine more than one unit, and a unit can appear in multiple-choice questions as well as in free response. The map below is a study tool, not a promise that a particular topic will appear in a particular numbered question.

AP unitMain focusWhat an exam prompt may ask you to do
Unit 1: Exploring One-Variable DataDescribe a single categorical or quantitative variableChoose or interpret a graph; compare distributions using context, shape, center, spread, or unusual features
Unit 2: Exploring Two-Variable DataDescribe relationships between two variablesInterpret an association, a two-way table, or a regression summary; assess whether a linear model is useful
Unit 3: Collecting DataSampling, surveys, observational studies, and experimentsEvaluate a design; identify bias or confounding; explain what conclusions the plan supports
Unit 4: Probability, Random Variables, and Probability DistributionsModel chance outcomes and their distributionsFind or interpret probabilities; use a suitable probability model; explain expected behavior in context
Unit 5: Sampling DistributionsDescribe how a statistic varies from sample to sampleReason about the behavior of a sample proportion or sample mean, including its center, spread, and shape

This naming matters. Regression as a descriptive method for two quantitative variables belongs with Unit 2’s two-variable data topics, not Unit 5. Unit 5 is Sampling Distributions. Inference for regression slopes is a later topic, in Unit 9; recognizing a regression line in a Unit 2 setting does not mean you should use a slope test or interval before that topic has been taught.

A Practical Map for Four Free-Response Questions

Think of the four free-response questions as opportunities to show different kinds of statistical reasoning, rather than as four fixed unit containers. A prompt may ask for a written explanation, a calculation, a design critique, or several connected steps. Read its verbs and context before deciding which ideas to use.

  • A data description or analysis task. Unit 1 may supply the tools for describing one variable. Unit 2 may be needed when the question concerns a relationship between two variables. A good response chooses evidence that answers the question, rather than listing every graph or statistic you know.
  • A data-collection or study-design task. Unit 3 is central when you must describe a sampling plan, identify how treatments were assigned, or decide whether a design supports a population description or a cause-and-effect conclusion. Unit 1 or Unit 2 may also matter because the response variable or relationship must be clearly identified.
  • A chance or probability-model task. Unit 4 supplies probability reasoning and models for random outcomes. Be clear about what the random variable represents and what event the probability describes. Some questions may ask you to interpret a result rather than calculate one.
  • A sampling-variation or synthesis task. Unit 5 asks how a statistic, such as \(\hat{p}\) or \(\bar{x}\), would vary across repeated samples under stated conditions. This reasoning is also a foundation for inference in later units. A synthesis prompt can combine sampling distributions with design details, data summaries, or a conclusion about evidence.

A useful planning routine is to mark the requested task, name the variable or statistic involved, identify the relevant unit ideas, and check what the question says about how the data were obtained. Then answer only what the prompt asks, with enough context to make each claim clear.

Exam-reading rule: Do not decide that a prompt is “a Unit 3 question” merely because it mentions an experiment, or “a Unit 5 question” merely because it uses a sample statistic. Identify the requested conclusion and the evidence needed to support it. A single prompt can draw on several units.

How Multiple-Choice Sets Use the Same Map

Multiple-choice questions may be independent, or several questions may share a short description, graph, table, or study scenario. When questions share information, read the full stimulus once for the population, individuals, variables, and data-collection method. Then treat each question as a separate task: one might ask about a graph, another about a design limitation, and another about probability.

Units 1 and 2 often call for matching a representation or statistic to the question. For a single quantitative variable, a prompt about shape or unusual values calls for a different kind of evidence than a prompt about typical value. For two quantitative variables, a question about direction and strength differs from one about the meaning of a fitted slope or the pattern in residuals. Use the distinctions developed in earlier tutorials, such as “Selecting Evidence for a Written Conclusion” and “Reading Full Regression Output Step by Step.”

Unit 3 questions often hinge on what was actually done. A random sample and random assignment are not interchangeable: one helps support generalizing to a population, while the other helps support cause-and-effect reasoning. Unit 4 questions require attention to the chance process and the event being asked about. Unit 5 questions ask about the distribution of a statistic across repeated samples, not just the distribution of individual observations in one sample.

For a set of questions, avoid assuming that every part uses the same method. A shared scenario can support several different questions, and a calculation in one part does not automatically answer a later interpretation question. Re-read the exact wording, especially phrases such as “for the population,” “among the participants,” “if repeated samples were taken,” or “does the treatment cause.”

Worked Example: Describe One Variable, Then Choose the Right Evidence

Worked Example: Battery Runtime Data

Situation. A fictional engineering class records the runtime, in hours, of six rechargeable batteries in a classroom test: 5.8, 6.1, 6.3, 6.3, 6.7, and 8.8. A prompt asks students to describe a typical runtime and note whether the observations show any unusual feature.

Map the prompt. There is one quantitative variable, runtime, so Unit 1 is the primary connection. The question asks about a typical value and an unusual feature, not about a relationship or a cause. An ordered list or dotplot would make the distribution visible; the median and the overall pattern can help describe a typical result and the high value.

Calculate a center. The ordered values are already listed. With six observations, the median is the average of the third and fourth values:

$$ \text{Median}=\frac{6.3+6.3}{2}=6.3\text{ hours}. $$

The mean is also easy to check from the data:

$$ \bar{x}=\frac{5.8+6.1+6.3+6.3+6.7+8.8}{6} =\frac{40.0}{6}\approx 6.67\text{ hours}. $$

Interpret in context. The median runtime in this test was 6.3 hours. The 8.8-hour observation is noticeably above the other five values, so a graph would help assess how it affects the distribution’s shape and mean. The mean is greater than the median, consistent with the high value pulling the average upward. This description is limited to the six batteries tested; it does not, by itself, establish what runtimes to expect from every battery of this type.

Exam connection. A multiple-choice item might ask which statistic is less affected by the high observation; a free-response prompt might ask for a contextual description. Both draw on Unit 1, but the response should match the particular task.

Worked Example: Separate Design Evidence From Data Description

Worked Example: A School Survey Plan

Situation. A fictional school has 900 students in grades 9–12. A student council wants to estimate the proportion of enrolled students who usually bring lunch from home. It randomly selects 25 students from each grade’s enrollment list and asks them the same question. Of the 100 selected students, 84 respond.

Map the prompt. The goal is to describe a population proportion, so Unit 3’s sampling methods are central. The outcome is a categorical variable—whether a student usually brings lunch from home—and the population is the school’s 900 enrolled students. Unit 1 ideas can help describe the categorical responses, but the sampling plan determines how far the description can reasonably extend.

Evaluate the plan. Selecting students at random from each grade is a stratified random sample, with grade as the stratum. It ensures that every grade is represented in the selected group. However, 16 of the 100 selected students did not respond. The council should consider whether students who did not respond might have different lunch habits from respondents. Random selection does not remove a possible nonresponse concern.

State what the result can support. If the sampling and response process is carried out as described, the responses can provide evidence about students enrolled at this school, subject to possible nonresponse bias and measurement error. The survey does not support a cause-and-effect claim because no treatment was assigned. It also does not directly represent students at other schools.

Exam connection. A free-response question might ask for the sampling method and a limitation; a multiple-choice set might ask which population the survey represents or what nonresponse could affect. In either format, name the design feature and explain its consequence rather than saying only that “the sample might be biased.”

Worked Example: Identify a Sampling Distribution Question

Worked Example: Repeated Samples of a Proportion

Situation. Suppose 40% of the students in a large school district use a particular transit pass. A question asks about the sample proportion \(\hat{p}\) among repeated random samples of 100 students. Assume the district has 2,000 students and that each sample is an SRS.

Map the prompt. The question is not asking for the distribution of individual students’ transit-pass use; it asks how the statistic \(\hat{p}\) varies from sample to sample. That is a Unit 5 sampling-distribution question, using the probability and normal-model ideas developed earlier.

Check conditions. The samples are described as random. The 10% condition holds because \(100\) is no more than 10% of \(2{,}000\). The Large Counts condition holds for the approximate normal model because \(np=100(0.40)=40\geq 10\) and \(n(1-p)=100(0.60)=60\geq 10\).

Find the center and spread. The mean of the sampling distribution of \(\hat{p}\) is \(p=0.40\). Its standard deviation is:

$$ \sigma_{\hat{p}}=\sqrt{\frac{p(1-p)}{n}} =\sqrt{\frac{(0.40)(0.60)}{100}} =\sqrt{0.0024}\approx 0.0490. $$

The sampling distribution is approximately normal under the stated conditions. To estimate the chance that a sample proportion is at least 0.50, standardize:

$$ z=\frac{0.50-0.40}{0.0490}\approx 2.04, \qquad P(\hat{p}\geq 0.50)\approx 0.0206. $$

Interpret in context. If many independent SRSs of 100 students were drawn under these conditions, about 2.06% of their sample proportions would be 0.50 or higher. This is a probability statement about the sampling process, not a claim that an individual student has a 2.06% chance of using the pass.

Exam connection. A multiple-choice question may ask for the standard deviation or the meaning of this probability. A free-response question may also ask you to justify the model by checking conditions. Unit 5 is the key connection; Unit 4 supports the probability calculation.

Common Mistakes and AP Exam Tips

  • Assigning one unit to each exam question. The questions can combine topics, and there is no reliable one-question-per-unit pattern. Identify the task and evidence required by the wording.
  • Misnaming Unit 5. Unit 5 is Sampling Distributions. Regression describing two-variable data belongs with Unit 2; inference for regression slopes is a later Unit 9 topic.
  • Confusing data distributions with sampling distributions. A data distribution describes observed individuals. A sampling distribution describes a statistic over repeated samples. State which one the prompt concerns.
  • Giving a design label without its implication. “Stratified sample” or “randomized experiment” is not a complete explanation. Say what was divided or assigned and what kind of conclusion the method can support.
  • Reporting a statistic without answering the question. A calculation earns its value when you interpret it in context and connect it to the requested claim. Include the population or sample, variable, and units where relevant.
  • Overclaiming from a scenario. A sample survey does not establish causation, and random assignment alone does not make participants representative of a wider population. Separate generalization from cause-and-effect reasoning.
  • Assuming every question in a multiple-choice set uses the same idea. Shared context does not mean shared task. Re-check what each item asks before choosing evidence or a method.

For a full-credit response, make the reasoning visible. Name the relevant variable or statistic, give the method or evidence the prompt requests, and interpret the result in the stated setting. If a condition is required, state how the scenario meets it. If the prompt asks what a design can support, connect the design feature directly to the scope of the conclusion.

Key takeaway: Use the unit map to recognize relevant tools, not to predict a fixed exam layout. Units 1–5 move from describing data and relationships, through collecting data and modeling chance, to understanding sampling variation. Across free-response questions and multiple-choice sets, answer the specific task, connect evidence to the claim, and keep the conclusion within the design’s limits.

Check Your Understanding

For each prompt, identify the main unit connection and the kind of reasoning the question calls for.

  1. A question gives a dotplot of commute times for 30 students and asks for a description of the distribution. Which unit is central, and what features should a contextual description address?
  2. A study assigns volunteers by chance to two different reminder schedules. Which part of the design supports a cause-and-effect conclusion, and what limitation remains for generalizing to all students?
  3. A question asks how the sample mean would vary across repeated random samples. Is it asking about a data distribution or a sampling distribution? Which unit is central?
  4. A multiple-choice set includes a survey description, a graph, and a question about nonresponse. Why should you evaluate each item’s task separately even though the items share a scenario?
  5. Where does descriptive regression fit in the unit map, and in which unit is inference for regression slopes addressed?