Tutorials › AP Statistics › Choosing Procedures From Summary Tables

Choosing a mean-inference procedure · Tutorial 776 of 1000

Choosing Procedures From Summary Tables

Use the table’s row structure, study design, and research question to choose the mean-inference procedure that matches the data.

Intermediate 9 min read

What You'll Learn

  • Distinguish paired columns from separate group columns by checking how observations were collected.
  • Match a one-population mean, a mean difference, or a difference between two means to its procedure.
  • Tell when summary statistics are enough to choose a procedure but not enough to check its conditions.
  • Recognize when separate before-and-after summaries cannot supply the information needed for paired t.
  • Connect the chosen t procedure to the question’s goal: estimating or testing.

Read the Design Behind the Table

In “Matching Procedures to Conditions,” you learned to check the study design and data shape for a chosen mean procedure. A summary table can help you make that choice, but its layout alone does not determine the procedure. First ask what each row represents, whether measurements are linked, and what population parameter the question targets.

A table might put two measurements in adjacent columns, such as “before” and “after,” or it might put two groups in separate columns. Those layouts can look similar while representing different designs. Before choosing a procedure, look for a row identifier or a description of how the observations were collected. The same person measured twice creates a pair; two different people in two groups do not become a pair merely because their values appear on the same row.

Key idea: Choose from the study design and target parameter, not from the number of columns. One sample compared with a fixed value calls for one-sample t; linked measurements call for paired t on differences; two independent groups call for unpooled two-sample t.

A Table-Reading Routine

Start with the quantitative response and the question. Is the goal to estimate or test a mean? Is the target one population mean, the mean of within-pair differences, or the difference between two population means? As in “Identifying the Parameter in a Mean Problem,” define that target before matching a procedure.

Next, use the table headings and the study description together. A useful table may list one row per individual, one row per matched pair, or one summary row per group. The sample sizes and means describe the data, but they do not reveal by themselves whether observations are linked. Design details—such as “the same participants” or “different randomly assigned participants”—settle that question.

1
Identify the response and goal.
Confirm that the response is quantitative, then determine whether the question asks for an estimate or a test.
2
Name the target parameter.
Choose \(\mu\), \(\mu_d\), or \(\mu_1-\mu_2\), according to the population quantity being asked about.
3
Trace the observational units.
Check whether the same units were measured twice, distinct units were deliberately matched, or the groups are independent.
4
Match the procedure and check what the table can show.
Choose the corresponding t procedure. Then check whether the design and relevant data-shape information are available; summary statistics alone may not establish every condition.

For a one-sample t procedure, the table should summarize one sample whose mean is compared with a fixed benchmark. For paired t, the analysis uses one difference per pair. A table of differences is especially useful because it can provide the sample size, mean difference, and standard deviation of differences. For an unpooled two-sample t procedure, the table summarizes two independent groups, usually with each group’s sample size, sample mean, and sample standard deviation.

A table with separate “before” and “after” summary rows may give each measurement’s mean and standard deviation but omit the standard deviation of the within-person differences. Those separate summaries are not enough to reconstruct the paired t statistic. The design still tells you that paired t is the matching procedure; the table may simply lack the information needed to carry out that analysis.

Worked Examples: Match the Table to the Target

Worked Example: A Paired Summary Table for Recovery Time

A fictional rehabilitation team measures how many minutes 16 randomly selected clients need to complete a routine before and after a program. The same clients are measured twice. The question is whether the program reduces mean completion time. The table summarizes the differences defined as after minus before:

Quantity for the differencesSummary
Number of clients, \(n\)16
Mean difference, \(\bar{x}_d\)\(-3.2\) minutes
Standard deviation of differences, \(s_d\)4.0 minutes
Shape of differencesApproximately symmetric; no apparent outliers

State: Let \(\mu_d\) be the true mean change in completion time, in minutes, for the population represented by the random sample, where \(d=\text{after}-\text{before}\). A reduction corresponds to a negative difference. Test \(H_0:\mu_d=0\) against \(H_a:\mu_d<0\).

Plan: Use a paired t test because the same clients contribute both measurements, and the table summarizes one difference per client. The 16 clients were randomly selected from 250, and \(16\leq0.10(250)=25\), so the 10% condition is satisfied. The differences from distinct clients are treated as independent. With 16 differences, the shape condition is supported by the approximately symmetric distribution and lack of apparent outliers.

Do: The standard error of the mean difference is \(s_d/\sqrt{n}=4.0/\sqrt{16}=4.0/4=1.0\) minute. The test statistic is

$$ t=\frac{\bar{x}_d-0}{s_d/\sqrt{n}} =\frac{-3.2-0}{4.0/\sqrt{16}} =\frac{-3.2}{1.0} =-3.20 $$

The degrees of freedom are \(n-1=16-1=15\). The left-tailed p-value for \(t=-3.20\) with 15 degrees of freedom is approximately \(0.0030\), rounded.

Conclude: Since \(0.0030<0.05\), reject \(H_0\). The data provide convincing evidence that the mean completion time after the program is lower than the mean completion time before the program for the population represented by the sampled clients.

This table supplies the difference statistics needed for paired t. If it had reported only separate before and after means and standard deviations, it would still identify the paired design, but it would not provide \(s_d\) for this calculation.

Worked Example: Two Independent Groups in Separate Columns

In a fictional randomized experiment, 40 students are assigned to one of two independent study routines, with 20 students in each group. Their response is the number of practice problems completed during a fixed session. Researchers want to estimate how much the population mean differs between the routines. The summary table is:

Study routine\(n\)\(\bar{x}\), problems\(s\), problems
Routine A2018.44.0
Routine B2021.05.0

Identify the target: Let \(\mu_A\) and \(\mu_B\) be the population mean numbers of problems completed under routines A and B. The target is \(\mu_A-\mu_B\), in problems. The request is to estimate that difference, so the goal is a confidence interval rather than a test.

Choose the procedure: Different students are assigned to the two routines, and each student contributes one response to one group. There are no repeated measurements or deliberate matches. Use an unpooled two-sample t interval for \(\mu_A-\mu_B\), as in “Experiments Comparing Two Treatments.”

Check what is known: Random assignment supports comparing the routines in this experiment, and each student contributes to only one independent group. The table provides \(n\), \(\bar{x}\), and \(s\) for each group, which are the summary statistics used by the procedure. The table alone does not show the individual response distributions, so it cannot establish the shape condition. The study description or graphs of the responses in each group are needed for that check. Because the students were assigned to treatments rather than described as a random sample of all students, the experiment does not by itself justify generalizing to every student.

The separate group columns are consistent with two-sample t here because the design confirms that the groups contain different, unlinked students. Had the columns listed the same students’ responses under both routines, the analysis would instead require paired differences.

Worked Example: One Sample Compared With a Benchmark

A fictional bottling line is checked by taking a random sample of 36 bottles. The response is fill volume in milliliters. A summary table reports \(n=36\), \(\bar{x}=502.4\) milliliters, and \(s=6.0\) milliliters. The quality question is whether the population mean differs from the target fill of 500 milliliters.

Identify the target: Let \(\mu\) be the true mean fill volume, in milliliters, for bottles from this line. The target is one population mean compared with the fixed benchmark of 500 milliliters—not a difference between two sample means.

Choose the procedure: Use a one-sample t test, with \(H_0:\mu=500\) and \(H_a:\mu\ne500\). If the question instead asked for a range of plausible values for the mean fill, the matching procedure would be a one-sample t interval.

Check what is known: The bottles were randomly sampled, and 36 is the sample size. The summary table does not state the size of the production population, so the 10% condition cannot be checked from the table alone. Nor do \(n\), \(\bar{x}\), and \(s\) reveal whether the individual fill volumes have strong skewness or outliers. A graph or more information about the measurements is needed for the shape check. Do not mistake the single sample’s summary row for evidence of two independent groups.

The procedure is selected by the target and design even when the table does not contain enough information to verify every condition. A careful response can name the matching procedure while also noting which checks remain unresolved.

When a Table Leaves Out the Crucial Link

Consider a table that reports the mean and standard deviation of “before” values and the mean and standard deviation of “after” values. If the same people were measured twice, the design is paired. However, separate summaries do not give the standard deviation of the differences. In particular, the standard deviation of before values and the standard deviation of after values cannot simply be subtracted to obtain \(s_d\). The relationship between each person’s two measurements affects the variability of the differences.

If the table instead reports two treatment groups, ask whether the groups consist of different units. When they do, and no deliberate matching is described, the target is the difference between two population means and the matching procedure is unpooled two-sample t. If the same units appear in both columns, or two units in each row were deliberately matched, the target is a mean of differences and the matching procedure is paired t.

Important distinction: A procedure can be identifiable even when a table does not provide enough detail to perform the analysis. Use the study description to choose the procedure; use the relevant data or graphs to check conditions and obtain any missing statistics.

This distinction also applies to one-sample summaries. A row containing \(n\), \(\bar{x}\), and \(s\) may be enough to calculate a one-sample t statistic or interval, but those three numbers do not tell you whether the data came from a suitable random process or whether a small sample contains an outlier. Summary statistics describe selected features of the data; they do not replace the design description or a shape check.

Common Mistakes and AP Exam Tips

  • Choosing by column count: Two columns do not automatically mean two independent samples. State whether the same units appear in both columns or whether the groups contain different units.
  • Assuming side-by-side rows are pairs: A shared row number creates a pair only if the study design deliberately links the observations. If rows merely line up for display, that is not a genuine pair.
  • Using paired t without differences: Paired t analyzes one difference per pair. Separate before-and-after standard deviations do not supply the standard deviation of those differences.
  • Using two-sample t on repeated measurements: Repeated measurements from the same individual are dependent. Analyze their differences rather than treating the columns as independent groups.
  • Confusing a benchmark with a second group: A fixed target such as 500 milliliters is not a second sample. One sample compared with a fixed value calls for one-sample t.
  • Claiming all conditions are met from summary statistics: Values of \(n\), \(\bar{x}\), and \(s\) do not establish random sampling, independence, or the shape of the data. Identify what the table shows and what information is still needed.
  • Ignoring the question’s goal: A request to estimate calls for an interval; a request to assess evidence for a claim calls for a test. The design identifies the kind of procedure, and the goal identifies interval versus test.

A full-credit choice names the target parameter, identifies how the observational units are connected, and states the matching procedure. For example: “Because the same sampled clients were measured before and after, define \(d=\text{after}-\text{before}\) and use paired t for \(\mu_d\).” If the table omits needed information, say so explicitly rather than inventing a condition check or treating missing summaries as available.

Key takeaway: Read a summary table through the study design: one sample versus a fixed value points to one-sample t, genuine pairs point to paired t on differences, and separate unlinked groups point to unpooled two-sample t. The table can guide the choice, but it may not contain enough information to verify conditions or complete the analysis.

Check Your Understanding

For each situation, identify the target parameter and the matching mean procedure. Note any important information the table alone does not establish.

  1. A table gives one sample’s \(n\), \(\bar{x}\), and \(s\), and the question compares its population mean with a fixed standard. Which procedure matches?
  2. Two columns list scores for the same 18 students before and after a workshop. What procedure matches, and which standard deviation is needed?
  3. A table gives separate summaries for two groups of different randomly assigned participants. What parameter and procedure match a request to estimate the mean difference?
  4. Separate before-and-after summaries give both means and standard deviations but no summary of individual differences. Why can you identify the design but not calculate the paired t statistic from those summaries alone?
  5. Two columns contain measurements displayed on the same rows, but the study says the participants in the groups are different and unlinked. Should the rows be treated as pairs? Explain.