Make Both Parameters Say What They Measure
In Hypotheses for Comparing Two Population Proportions, you practiced writing hypotheses with \(p_1\) and \(p_2\). Before those hypotheses can be interpreted, each parameter needs a clear definition. A reader should be able to tell which group each proportion describes, what outcome counts as success, and—in a study—what population the group represents.
For example, “\(p_1\) is the proportion in Group 1” is not enough. Which people or units are in Group 1? Proportion of them with what characteristic? A strong definition answers both questions, then defines \(p_2\) in a parallel way for the second group using the same outcome.
Keep the distinction between parameters and sample statistics from earlier in this course. The parameters \(p_1\) and \(p_2\) describe the populations or conditions of interest. The sample proportions \(\hat{p}_1\) and \(\hat{p}_2\) describe the observed samples. A definition should not use sample counts or say that \(p_1\) is the proportion “in the survey.”
A Reliable Pattern for Defining Two Proportions
Use the following pattern when writing definitions:
Identify the population, naturally occurring group, or experimental condition represented by the parameter.
Describe exactly what counts as success, using the same outcome for both parameters.
Say “the true proportion” and identify the relevant population or experimental units, rather than describing only the observed sample.
Say which group is \(p_1\) and which is \(p_2\), so that \(p_1-p_2\) has a clear interpretation.
A useful structure is: “Let \(p_1\) be the true proportion of [clearly named group] that [has the shared outcome]. Let \(p_2\) be the true proportion of [second group] that [has the same outcome].” Adapt the group wording to the study design. In an observational study, groups are often populations or naturally occurring categories. In an experiment, they are defined by the treatment conditions.
Observational Studies: Name the Populations
In an observational study, researchers record characteristics or outcomes without assigning people to treatments. The groups may be different places, age categories, or other naturally occurring groups. Define each parameter as a true proportion in the population that the group is meant to represent—not merely among the people who happened to respond.
Be specific about the target population. “People who recycle” could mean residents of a particular city, households in a region, or another population. If the question concerns a sample drawn from two communities, the parameter definitions should name those communities’ resident populations and the characteristic being measured.
Worked Example: Comparing Two Communities in a Survey
Question: A fictional environmental researcher surveys separate samples of residents in Pineview and Lakefield. Each resident is asked whether their household composted food scraps at least once during the past month. Define \(p_1\) for Pineview and \(p_2\) for Lakefield.
Identify the groups and outcome: The groups are the populations of residents in the two named communities. Success is the same for both groups: the resident’s household composted food scraps at least once during the past month.
Define the parameters: Let \(p_1\) be the true proportion of all Pineview residents whose household composted food scraps at least once during the past month. Let \(p_2\) be the true proportion of all Lakefield residents whose household composted food scraps at least once during the past month. The comparison order is Pineview minus Lakefield.
These definitions do not say that \(p_1\) is the proportion among the surveyed Pineview residents. The survey produces a sample proportion, \(\hat{p}_1\); \(p_1\) describes the population proportion the survey is intended to estimate. The same distinction applies to \(\hat{p}_2\) and \(p_2\).
The wording of the outcome matters. “Composts” is vague: it could mean ever, regularly, or during a particular period. Naming the exact recorded outcome makes the two parameters comparable. If one definition says “composted last month” and the other says “has ever composted,” the parameters do not refer to the same characteristic.
Worked Example: Defining Proportions for Two Age Groups
Question: A fictional public-health survey studies adults in a county. Participants report whether they received a seasonal flu vaccine during the most recent flu season. The comparison is between adults ages 18–39 and adults ages 40–64. Define \(p_1\) for the younger group and \(p_2\) for the older group.
Identify the groups and outcome: The two groups are county adults in the specified age ranges. Success is having received a seasonal flu vaccine during the most recent flu season.
Define the parameters: Let \(p_1\) be the true proportion of all adults ages 18–39 in the county who received a seasonal flu vaccine during the most recent flu season. Let \(p_2\) be the true proportion of all adults ages 40–64 in the county who received a seasonal flu vaccine during that same flu season. The group order is ages 18–39 minus ages 40–64.
The shared time period is part of the outcome definition, not an optional detail. Each parameter concerns the same county, the same kind of event, and the same flu season; only the age group changes. Defining the groups this way also avoids treating an individual survey respondent as the population.
Experiments: Name the Treatment Conditions
In an experiment, researchers assign experimental units to treatment conditions. The groups are therefore not naturally occurring populations such as Pineview and Lakefield. They are the conditions being compared, such as a new training program and a standard program. Define each parameter by naming the condition and the outcome measured for units assigned to it.
A clear definition identifies the experimental units and the success outcome, then distinguishes the treatment condition from the comparison condition. For example, “\(p_T\) is the true proportion of experimental units that show the outcome under the new treatment” is more informative than “\(p_T\) is the treatment group’s proportion.” The shorter version leaves the units and outcome unspecified.
The order is often treatment minus control, written \(p_T-p_C\), but the symbols must be defined rather than assumed. In the examples below, \(p_T\) refers to the treatment condition and \(p_C\) to the control condition. The definitions do not by themselves establish a treatment effect; they state which proportions are being compared. Earlier tutorials explain how random assignment affects the scope of a causal conclusion.
Worked Example: Comparing Two Study-Support Methods
Question: In a fictional experiment, high-school students are randomly assigned either to use a text-message study reminder program or to receive the school’s usual study-support materials. The outcome is whether a student completes at least four planned study sessions during a two-week period. Define the two proportions.
Identify the experimental units, conditions, and outcome: The experimental units are the high-school students in the experiment. The treatment condition is use of the text-message reminder program, and the control condition is receipt of the usual study-support materials. Success is completing at least four planned study sessions during the two-week period.
Define the parameters: Let \(p_T\) be the true proportion of experimental units who would complete at least four planned study sessions during the two-week period under the text-message reminder program. Let \(p_C\) be the true proportion of experimental units who would complete at least four planned study sessions during the same period under the usual study-support materials. The order is treatment minus control.
The two definitions use the same units and outcome while changing the assigned condition. They do not define \(p_T\) as the proportion of students who happened to receive messages and \(p_C\) as the proportion who did not, without explaining what the conditions and success outcome are. Those details are what make the parameters interpretable.
Worked Example: Naming the Outcome in a Medical Experiment
Question: A fictional experiment assigns volunteers with seasonal allergies to a new nasal spray or a comparison spray. After seven days, researchers record whether each volunteer reports improved symptoms. Define \(p_T\) and \(p_C\), with the new spray as the treatment.
Identify the units, conditions, and outcome: The experimental units are the volunteers in the experiment. The treatment is the new nasal spray, the comparison condition is the comparison spray, and success means reporting improved allergy symptoms after seven days.
Define the parameters: Let \(p_T\) be the true proportion of experimental units who would report improved allergy symptoms after seven days under the new nasal spray. Let \(p_C\) be the true proportion of experimental units who would report improved allergy symptoms after seven days under the comparison spray. The order is new spray minus comparison spray.
“Improved symptoms” is the shared outcome, and “after seven days” fixes when it is assessed. If the response were instead defined as having no symptoms, the parameters would describe a different outcome. Do not quietly change the success criterion between the treatment and comparison definitions.
Connect Definitions to the Hypotheses
Once the parameters are defined, they give meaning to the hypotheses. If the group order is Pineview minus Lakefield, for instance, \(p_1-p_2\) means the true composting proportion in Pineview minus the true composting proportion in Lakefield. If the order is treatment minus control, \(p_T-p_C\) means the treatment-condition proportion minus the control-condition proportion.
The hypotheses must refer to the same parameters and outcome as the definitions. As in Hypotheses for Comparing Two Population Proportions, the standard null for no difference states equality, while the alternative matches the research question. Writing the definitions first helps prevent a symbol or inequality from being attached to the wrong group.
Worked Example: Check the Definitions Against a Directional Claim
Question: In a fictional experiment, students are assigned to either a new online review activity or the usual review activity. Success is earning at least 80 percent on a quiz one week later. The question asks whether the new activity increases the proportion meeting that benchmark. Define the parameters and write hypotheses consistent with the stated order.
Define the parameters: Let \(p_T\) be the true proportion of experimental units who would earn at least 80 percent on the quiz one week later under the new online review activity. Let \(p_C\) be the true proportion of experimental units who would earn at least 80 percent on the quiz one week later under the usual review activity.
Match the hypotheses to the definitions: “Increases” means the new-activity proportion is greater than the usual-activity proportion. With treatment first, the null and alternative are:
The alternative matches the parameter order and the research question. If the definitions were instead written with the usual activity first, the same claim would require the opposite inequality. The wording of a claim does not change; the sign changes when the group order changes.
Common Mistakes and AP Exam Tips
- Defining a sample statistic instead of a parameter: “The proportion among the 80 surveyed residents” describes an observed sample proportion. Define \(p_1\) and \(p_2\) for the target populations or experimental conditions.
- Leaving the group unnamed: “\(p_1\) is the proportion who responded yes” does not identify which population or condition \(p_1\) represents. Name that group explicitly.
- Using different success outcomes: Both parameters must refer to the same characteristic, with the same time frame or measurement rule when relevant.
- Calling experiment groups populations without explanation: In an experiment, name the experimental units and the assigned conditions. Do not confuse treatment assignment with naturally occurring membership in a population.
- Using “treatment proportion” without details: State which treatment and what outcome is measured. A label alone does not define the parameter.
- Changing the order halfway through: Keep the same order in the definitions, difference, hypotheses, and written explanation. If Group 1 is the new program, \(p_1-p_2\) is new program minus comparison.
- Writing only “true proportion”: That phrase signals a parameter, but it is incomplete unless the group and success outcome are also clear.
Key Takeaway
Good parameter definitions make a two-proportion comparison understandable before any calculation begins. Observational studies call for clear population or naturally occurring group wording; experiments call for clear experimental-unit and treatment-condition wording. In either setting, define the same success outcome for both proportions and keep the group order fixed.
Check Your Understanding
For each situation, practice defining both parameters precisely. Keep the outcome and group order clear.
- A fictional survey compares residents of two towns on whether they used public transit at least once in the past week. What information should the definitions of \(p_1\) and \(p_2\) include?
- A study compares the proportion of adults ages 25–44 and ages 45–64 in a county who have a current library card. Define \(p_1\) for the younger group and \(p_2\) for the older group.
- In an experiment, volunteers receive either a new reminder app or the usual paper reminder. Success is attending a scheduled appointment. What details belong in definitions of \(p_T\) and \(p_C\)?
- Why is “\(p_1\) is the proportion of people who answered yes” an incomplete parameter definition?
- An experiment compares a new gardening lesson with the usual lesson. The outcome is planting at least one native plant within a month. Write definitions for \(p_T\) and \(p_C\), with the new lesson first.