What Can Three Summary Statistics Tell You?
A report may give only the sample size \(n\), sample mean \(\bar{x}\), and sample standard deviation \(s\). Those numbers are enough to calculate a one-sample t interval or test, but they do not automatically show that the procedure is appropriate. Before using mean inference, you still need evidence about how the observations were collected, whether they can be treated as independent, and—especially for a small sample—whether the data’s shape supports the t procedure.
In “Verifying Conditions From a Described Study,” we used a condition evidence audit to separate conditions that are supported, not met, or unclear. Here, the challenge is that the information may be limited to summary statistics. Your goal is not to guess what the study did or what its data looked like. It is to identify which checks those summaries allow, make only justified claims, and name the additional information needed.
Audit the Evidence—Do Not Fill in the Gaps
For a one-sample t procedure, the condition audit still includes randomness, independence (including the 10% condition when relevant), and the Normal/Large Sample condition. The earlier tutorials on these checks explain the conditions themselves. The important question here is which parts can be assessed from \(n\), \(\bar{x}\), and \(s\), and which require information beyond the summaries.
| Condition or question | What the summaries can tell you | What additional information may be needed |
|---|---|---|
| Randomness | Nothing about how the observations were selected or assigned. | A description of the chance process and the population or study units it involved. |
| Independence | The sample size alone does not show whether observations are linked, repeated, or clustered. | How many observations came from each unit, whether units were linked, and relevant sampling details. |
| 10% condition | The summaries give \(n\), but not necessarily the population size \(N\). | For sampling without replacement, the population size, or other information that lets you check \(n\leq0.10N\). |
| Normal/Large Sample | The sample size shows whether \(n\geq30\). Neither \(\bar{x}\) nor \(s\) shows the distribution’s shape. | For \(n<30\), a graph of the observations or information that the population distribution is approximately Normal. |
The Normal/Large Sample condition has an important distinction. As established in “Checking the Normal/Large Sample Condition” and “Using the n at Least 30 Rule Correctly,” \(n\geq30\) supports the large-sample route. If \(n<30\), the summaries alone cannot reveal whether the data are roughly symmetric, strongly skewed, or affected by outliers. A mean and standard deviation do not preserve enough information to reconstruct a dotplot, histogram, boxplot, or Normal probability plot.
A small sample is not automatically unsuitable. It means you need shape evidence that the summaries do not provide. Conversely, \(n\geq30\) helps with the shape check but does not establish that the sample was random or that observations are independent. Conditions are separate; evidence for one cannot fill a gap in another.
A Practical Workflow for Summary-Only Questions
Determine whether the question concerns one mean, paired differences, or a difference between two means. Summary statistics do not tell you how the groups or observations were formed.
For each relevant sample or set of differences, compare \(n\) with 30. If \(n\geq30\), the large-sample route supports the Normal/Large Sample condition. If \(n<30\), request shape evidence.
Ask whether units were randomly sampled or randomly assigned, whether observations are repeated or linked, and—when sampling without replacement—what population size applies.
For every condition, identify what is supported, what is not met, or what is unclear. Then state whether the available evidence justifies proceeding with the intended inference.
The same logic applies when two groups are summarized separately. Check the sample size for each group, and do not use one group’s shape evidence to stand in for the other’s. For paired data, the relevant quantitative observations for a paired t procedure are the pairwise differences; summaries of two groups separately do not reveal whether observations were paired or what the differences look like.
Worked Examples
Worked Example: A Small Sample With Only Summary Statistics
A fictional community garden report gives \(n=18\), \(\bar{x}=6.4\) kilograms, and \(s=1.1\) kilograms for the weekly harvest from each of 18 garden plots. The report says nothing about how the plots were selected, whether they were selected from a larger list, or what the data look like. The intended analysis is a one-sample t procedure for the mean weekly harvest of plots in the garden network.
State. Let \(\mu\) be the mean weekly harvest, in kilograms, for plots in the garden network. We need to determine whether the reported summaries verify the conditions for a one-sample t procedure and what information is missing.
Plan. Audit randomness, independence, and the Normal/Large Sample condition. The summaries include \(n\), so we can assess the sample-size route. Because the sample was described as coming from plots, we also need the selection method and details about the plots to assess randomness and independence. If sampling was without replacement, we need the population size to check the 10% condition.
Do. The value \(n=18\) is less than 30, so the large-sample route is not met. The report gives no graph and no information that the population distribution is approximately Normal, so the Normal/Large Sample condition is unclear. The values of \(\bar{x}=6.4\) and \(s=1.1\) do not establish the data’s shape. The report also does not describe random selection or assignment, so the Random Condition is unclear. It does not say whether observations are independent or provide a population size for a 10% check; those independence details are also unclear.
Conclude. The summaries alone do not verify the conditions, so we cannot say that a one-sample t procedure is justified from this report. We would need the plots’ selection method and relevant independence details, the population size if sampling without replacement, and a graph of the 18 harvest values or information about the population’s shape.
Worked Example: A Large Sample With Missing Design Information
A fictional device-testing report gives \(n=46\), \(\bar{x}=312\) hours, and \(s=38\) hours for battery life. It does not say how the devices were chosen, whether multiple batteries came from the same device batch, or what population of batteries the report is meant to represent.
State. Let \(\mu\) be the mean battery life, in hours, for the population the testing team intends to describe. We will assess what the summaries establish and what study information is still needed.
Plan. Check the Normal/Large Sample condition using \(n\), but audit randomness and independence from the study design rather than from the numerical summaries. If the intended inference is based on sampling without replacement from a known finite population, the 10% condition requires the population size.
Do. Since \(n=46\geq30\), the large-sample route supports the Normal/Large Sample condition. However, no selection method is reported, so the Random Condition is unclear. The summaries do not tell us whether the 46 batteries represent independent units or whether they share batches or other sources of linkage. If they were randomly sampled without replacement, we would also need the population size \(N\) to check whether \(46\leq0.10N\). The mean and standard deviation provide no answer to these design questions.
Conclude. The sample size supports the Normal/Large Sample condition, but the report does not provide enough information to assess randomness or independence. Before using a one-sample t procedure to generalize to a population, we need to know how the batteries were selected, what population is represented, and whether the observations can reasonably be treated as independent.
Worked Example: Two Groups With Different Sample Sizes
A fictional recreation center compares weekly exercise time for two independently sampled groups of members. The report gives the following summaries: the morning group has \(n_1=34\), \(\bar{x}_1=142\) minutes, and \(s_1=36\) minutes; the evening group has \(n_2=22\), \(\bar{x}_2=128\) minutes, and \(s_2=31\) minutes. It states that each group was randomly sampled from its own membership list, but does not give either list’s size or any graphs.
State. Let \(\mu_1\) and \(\mu_2\) be the mean weekly exercise times, in minutes, for members represented by the morning and evening lists. We will audit the conditions for a two-sample t procedure for \(\mu_1-\mu_2\).
Plan. Assess randomness and independence for each sample and between the groups. Check the Normal/Large Sample condition separately for each group. If sampling without replacement, the 10% condition must be checked against each group’s membership-list size.
Do. Random sampling is stated for both groups, supporting the Random Condition for the populations represented by their respective lists. The sizes of the lists are not given, so we cannot check whether \(34\leq0.10N_1\) or \(22\leq0.10N_2\); the 10% checks are unclear. The description calls the samples independent, but to verify independence between groups we would want to confirm that members are not included in both groups or otherwise linked. For shape, \(n_1=34\geq30\), so the large-sample route supports the condition for the morning group. For the evening group, \(n_2=22<30\); without a graph or population-shape information, its Normal/Large Sample condition is unclear. The reported means and standard deviations do not resolve that gap.
Conclude. The summaries and description support random sampling and the large-sample route for the morning group, but do not establish every condition for a two-sample t procedure. We need membership-list sizes to check the 10% conditions, confirmation that the groups are independent, and shape evidence for the evening group.
Worked Example: Summary Statistics for Differences Do Not Describe the Pairing
A fictional physical therapy report summarizes change in a balance score for 16 clients, with \(\bar{x}_d=2.8\) points and \(s_d=1.5\) points. The report does not explain whether each value is a before-and-after difference for the same client, nor does it provide a graph or the individual differences.
State. If the 16 values are paired differences, let \(\mu_d\) be the mean change in balance score for the population represented by the clients. We need to decide whether these summaries alone justify a paired t procedure.
Plan. For paired t inference, first verify that the observations are genuine within-client differences and that different clients’ differences can be treated as independent. Then assess randomness and the Normal/Large Sample condition using the differences, not the two measurements separately.
Do. The report does not state whether the observations are paired differences, so the intended paired procedure cannot yet be confirmed. If \(n=16\) does refer to 16 client-level differences, then \(16<30\), and the summaries do not show whether those differences are reasonably symmetric or have pronounced outliers. We would need the study design and a graph of the differences or information about their population shape. We would also need to know how clients were selected or assigned and whether the differences from different clients are independent.
Conclude. The values \(n=16\), \(\bar{x}_d=2.8\), and \(s_d=1.5\) are not enough to verify the paired t conditions. We need confirmation of the pairing and study design, plus shape evidence for the individual differences because the sample is small.
Common Mistakes and AP Exam Tips
- Treating a reported mean and standard deviation as a shape description. A mean and standard deviation do not reveal skewness, clusters, gaps, or outliers. For a sample smaller than 30, request a graph or information about the population distribution.
- Claiming that a large \(n\) verifies every condition. The large-sample route supports the Normal/Large Sample condition. It does not show that selection was random, that observations are independent, or that the sample represents the target population.
- Assuming the 10% condition without \(N\). If the sample was taken without replacement and the population size is not given, say the condition cannot be checked from the available information. Do not simply assume the population is large.
- Ignoring the study design because the question gives numbers. A calculator can use \(n\), \(\bar{x}\), and \(s\), but those inputs do not establish how observations were obtained. Ask what one observation represents and how units entered the study.
- Using pooled summaries to assess each group’s shape. In a two-sample t procedure, check the Normal/Large Sample condition separately for each group. One group’s sample size or graph does not settle the other group’s condition.
- Writing “conditions are met” when evidence is absent. Full-credit wording distinguishes “supported,” “not met,” and “unclear.” Name the missing information rather than making an unsupported guess.
A strong AP response connects each status to evidence: “Because \(n=18<30\) and the report gives no graph or population-shape information, the Normal/Large Sample condition is unclear. The report also does not describe random selection, so the Random Condition cannot be verified. We would need the selection method and a graph of the observations or information about the population distribution.” This makes clear what the summaries establish and what they do not.
Check Your Understanding
For each situation, identify what the summaries support and what additional information is needed to assess mean-inference conditions.
- A report gives \(n=12\), \(\bar{x}=9.3\), and \(s=2.1\). What can you conclude about the Normal/Large Sample condition, and what shape information would help?
- A study reports \(n=55\), a sample mean, and a sample standard deviation. Does \(n\geq30\) establish random sampling? Explain.
- A random sample of 24 households is taken without replacement, but the report gives no population size. What is needed to check the 10% condition?
- Two independent groups have sample sizes 31 and 20, with means and standard deviations reported for both. Which group has the large-sample route, and what additional shape evidence is needed?
- A report gives summary statistics for 14 supposed before-and-after differences. What design information must be confirmed before treating them as paired data?