Statistics as a Connected Process
A statistical investigation is more than a calculation. It begins with a question that can be answered using data, continues with a plan for obtaining relevant data, and uses analysis to produce a conclusion. In “Unit 5 Regression Synthesis Review,” you connected graphs and numerical summaries to a cautious conclusion about a regression. The same habit applies across statistics: make each step serve the question, and make the conclusion fit the data and the way they were collected.
The four practices are Formulate Questions, Collect Data, Analyze Data, and Interpret Results. They form a useful sequence, but they also inform one another. If analysis reveals that a variable was recorded unclearly, for example, the question or data-collection plan may need revision.
A statistical question anticipates variability in the answers. “Which lunch will one particular student choose?” asks about one individual. “Which lunch would students at this school choose?” asks about a group whose responses may differ, so collecting data can help describe the pattern.
A Four-Practice Planning Route
Before calculating, write down the investigation’s target and design. Use the following route to check that the question, data, analysis, and conclusion match.
Specify the population or process, the individuals being studied, the variable or variables, and the specific feature you want to learn about.
Explain how observations will be obtained. Consider whether the method fits the question and whether selection, nonresponse, or measurement could create a problem.
Use graphs and summaries that suit the variable types and the question. Keep calculations tied to the observations actually collected.
Answer the original question in context. State what the design allows you to conclude, and avoid unsupported generalizations or causal claims.
Worked Example: School Lunch Choices
Worked Example: Which Lunch Would Students Choose?
Situation. A fictional high school is considering four entrees for a special lunch day: pasta, tacos, a vegetable bowl, and a chicken sandwich. The school wants to learn which option students would select if all four were offered.
Formulate Questions. The population is all currently enrolled students in grades 9–12 at the school. The individuals are those students, and the variable is each student’s selected entree from the four choices. A focused statistical question is: “What is the distribution of students’ stated choices among these four entrees?” This wording is more useful than asking whether students “like the menu,” because it identifies what response will be recorded and compared.
Collect Data. The school selects 80 students at random from the enrollment roster and asks each the same question: “If you had to choose one of these four entrees for the special lunch, which would you select?” Each student gives one response, and all 80 selected students respond. An anonymous response form can reduce pressure to give a socially preferred answer. Using the roster gives enrolled students a chance to be selected; asking only students who happen to visit the cafeteria could leave out students who bring lunch.
The plan also has limits. Students are choosing hypothetically, not placing actual orders, and they must choose from only the four listed options. Their answers may not perfectly predict what they will eat on the day. These details matter when interpreting the results.
Analyze Data. The recorded counts are 35 for pasta, 27 for tacos, 12 for the vegetable bowl, and 6 for the chicken sandwich. First check that all responses are accounted for:
For each entree, divide its count by the sample size of 80 to find its sample proportion. For pasta, for example:
The full summary is:
| Entree | Count | Sample proportion |
|---|---|---|
| Pasta | 35 | 43.75% |
| Tacos | 27 | 33.75% |
| Vegetable bowl | 12 | 15.00% |
| Chicken sandwich | 6 | 7.50% |
The proportions sum to \(100\%\), as expected. A bar graph of the four counts or proportions would make the category comparison visible. Pasta is the most frequently selected option in this sample; the category with the greatest count is the sample’s mode.
Interpret Results. Among the 80 randomly selected students who responded, pasta was the most common stated choice: 35 students, or 43.75%, selected it. Tacos were next, at 33.75%. The school can use these sample results as information about enrolled students’ stated choices, while remembering that a random sample does not guarantee an exact match to the population. The responses do not establish what students will actually eat, why they prefer an option, or what students at other schools would choose.
Check the connection. The question asks about a distribution of categories, so counts, proportions, and a bar graph are appropriate. If the school instead asked how satisfied students were on a rating scale, the variable and useful summaries would change. Choosing the analysis only after seeing which calculations are convenient would reverse the logic of the investigation.
Worked Example: Transportation to a Community Park
Worked Example: Describe How Adults Reach the Park
Situation. A fictional community district wants to describe how adult residents usually travel to its central park. A random sample of 60 adults is selected from a resident list. Each person chooses one response: walk, bicycle, bus, or car. All 60 sampled adults respond. The counts are 27 walking, 9 bicycling, 3 taking the bus, and 21 traveling by car.
Formulate Questions. The population is adult residents in the district, the individuals are those adults, and the variable is usual travel method to the central park. The question is: “What are the proportions of adult residents in the district who usually use each of these four travel methods?”
Collect Data. The random sample is designed to represent adults on the resident list, rather than only people seen at the park. That distinction is important: surveying park visitors would miss adults who do not go there and could overrepresent people who use particular travel methods. The method still depends on the list covering the intended population and on respondents understanding “usually” in a similar way.
Analyze Data. The counts total \(27+9+3+21=60\). Divide each category count by 60. For walking, \(27/60=0.45\), or \(45\%\). The proportions for bicycling, bus, and car are \(9/60=15\%\), \(3/60=5\%\), and \(21/60=35\%\), respectively. A bar graph would compare the four categories without suggesting that the categories have a meaningful numerical order.
Interpret Results. In this sample, walking is the most commonly reported usual method, at 45%, followed by car at 35%, bicycling at 15%, and bus at 5%. These results describe the sampled adults and provide evidence about the district’s adult residents, subject to the coverage and response limitations. They do not show that a particular transportation option causes park visits or explain why residents choose each method.
This example uses the same four practices as the lunch investigation, but the statistical question is about travel methods rather than menu choices. The variable is categorical in both cases, so comparing category counts or proportions is useful. The context determines what those categories mean and what limitations to mention.
Worked Example: Background Music and Task Time
Worked Example: Compare Two Assigned Conditions
Situation. A fictional teacher recruits 12 volunteers to complete the same short worksheet. Six volunteers are randomly assigned to work with quiet background music, and six are randomly assigned to work in silence. Completion time, in minutes, is recorded. The music group’s times are 18, 20, 17, 21, 19, and 17 minutes. The silence group’s times are 22, 24, 20, 23, 21, and 22 minutes.
Formulate Questions. The question is: “For these volunteers doing this worksheet, how do the completion times compare between the music and silence conditions?” The explanatory variable is assigned condition, with categories music and silence; the response variable is completion time in minutes. Because the question compares two groups on a quantitative response, comparing group means is a relevant descriptive analysis.
Collect Data. The students volunteered, and the teacher randomly assigned them to the two conditions. Random assignment helps make the groups comparable for the experiment, while the volunteer recruitment limits how confidently the results can be generalized to all students. The teacher should keep other conditions, such as the worksheet and instructions, the same for both groups so the assigned condition is the intended difference.
Analyze Data. Find the mean time in each group. The music times sum to \(18+20+17+21+19+17=112\), so their mean is \(112/6\approx18.67\) minutes. The silence times sum to \(22+24+20+23+21+22=132\), so their mean is \(132/6=22\) minutes. Defining the comparison as silence mean minus music mean gives:
In these data, the music group’s mean completion time is about 3.33 minutes lower than the silence group’s mean. A graph of the individual times would add useful information about the spread and overlap; the mean difference alone does not describe every student’s experience.
Interpret Results. The observed volunteers completed the worksheet faster on average in the music condition than in silence. Since condition was randomly assigned, the experiment’s design can support a cautious discussion of whether the assigned condition affected these participants’ times. The small group sizes mean that chance differences between groups are possible, and no significance test has been conducted here. Because participants volunteered, the results also do not automatically generalize to all students, other tasks, or different music.
This example shows why the collection step affects interpretation. A sample survey can describe responses and, when appropriately sampled, support conclusions about a population. A randomized experiment can help investigate a cause-and-effect question. Neither design licenses every possible conclusion; the question, analysis, and scope still have to align.
Common Mistakes and AP Exam Tips
- Leaving the question vague. “What do students think?” does not identify the population, variable, or result of interest. A stronger question names the group and what will be measured or classified.
- Confusing the sample with the population. Say “among the 80 sampled students” when describing the observed lunch choices. If you discuss all enrolled students, explain that the random sample provides evidence about them rather than treating the sample percentages as exact population values.
- Ignoring how people were selected. A large number of responses does not automatically make a sample representative. Full-credit reasoning identifies a relevant issue, such as undercoverage or nonresponse, and explains how it could affect the conclusion.
- Choosing a summary that does not fit the variable. Lunch entree is categorical, so a mean of entree labels is not meaningful. Counts, proportions, and a bar graph directly compare the categories.
- Reporting a calculation without context. “43.75%” is incomplete. State that 43.75% of the 80 sampled students selected pasta, and clarify what that sample result can tell the school.
- Turning association or a group difference into a cause claim. Random selection helps with representation; random assignment helps with cause-and-effect reasoning. They serve different purposes. Identify which was used before making a claim.
- Claiming a result is established by a small observed difference. The worksheet example gives a descriptive mean difference, not a significance test. Say what the observed data show and do not imply that chance variation has been ruled out.
- Forgetting the original question at the conclusion. A strong interpretation answers the question that was formulated, then adds a relevant limitation. It does not introduce a broader claim just because the calculations are finished.
For an AP response, make the four practices visible in your reasoning. Name the population and variable, describe how data were obtained, give analysis that addresses the question, and state a conclusion in context. When a limitation matters, explain its consequence: for instance, a hypothetical lunch choice may not predict an actual order, and volunteers may not represent all students.
Check Your Understanding
Use the four practices to plan or evaluate each response.
- In the lunch example, identify the population, individuals, and variable. Why does the question specify the four entrees?
- Calculate the sample proportion of students who selected tacos. State what that proportion describes.
- Give one reason the lunch-choice results might not perfectly predict actual orders on the special lunch day.
- In the park example, explain why sampling only people already at the park could affect the results.
- In the worksheet example, which part of the design supports a cause-and-effect interpretation, and which part limits generalization?