Tutorials › AP Statistics › Why Right Skew Pulls the Mean Above the Median

Describing quantitative distributions · Tutorial 84 of 1000

Why Right Skew Pulls the Mean Above the Median

Use numerical examples to see why high values in a right-skewed distribution tend to pull the mean above the median, while the median is less affected.

Beginner 9 min read

What You'll Learn

  • Explain why high values in a right tail can raise the mean.
  • Calculate and compare the mean and median for small data sets.
  • Describe how changing a high value can affect the mean and median differently.
  • Compare typical mean–median patterns for right-skewed, symmetric, and left-skewed distributions.
  • Avoid using the mean and median alone as proof of a distribution’s shape.

When a Distribution Has a Long Right Tail

In Describing Shape: Symmetric, Skewed, Uniform, you learned that a distribution is named for the direction of its tail. A right-skewed distribution has a longer tail toward larger values. In this tutorial, we connect that visible shape to two ways of describing center: the mean and the median.

The mean is the arithmetic average: add all the observations and divide by their number. The median is the middle value after the observations are put in order, or the average of the two middle values when there are an even number of observations. Both describe center, but they respond differently to extreme values. A few observations far out in a right tail can raise the mean considerably, while they often have little effect on the median.

Key idea: In many right-skewed distributions, the mean is greater than the median because the larger values in the right tail raise the arithmetic average. This is a useful pattern, not a rule that identifies shape by itself: examine the graph and context as well as the two measures of center.

The mean uses the value of every observation. When a high value is added, or when an existing observation becomes much larger, the total sum increases. That raises the mean. The median depends on the ordered position of the middle observation or observations. A high value at the end of the list might change the median very little or not at all.

This difference helps explain why the mean can sit to the right of the median on the number line in a right-skewed distribution. It does not mean that most observations are above the median. Rather, a relatively small number of large observations can pull the average toward the tail. The examples below show how the calculation works.

A Right Tail Can Pull the Mean Up

Worked Example: Waiting Times at a Service Desk

A fictional school records the waiting times, in minutes, of nine students visiting a service desk. The ordered times are \(12, 13, 14, 14, 15, 16, 17, 18,\) and \(41\). The last wait is much longer than the others, giving the list a high-end tail.

Plan. Find the mean by adding all nine times and dividing by nine. Since there are nine ordered observations, find the median at position \((9+1)/2=5\). Then compare the measures in the context of the waiting times.

Do. The total waiting time is \(12+13+14+14+15+16+17+18+41=160\) minutes. Therefore,

$$ \bar{x}=\frac{160}{9}\approx 17.78\text{ minutes} $$

The fifth ordered observation is 15 minutes, so the median is 15 minutes. The mean is about \(17.78-15=2.78\) minutes greater than the median.

Conclude. The mean waiting time is about 17.78 minutes, while the median waiting time is 15 minutes. The much longer wait of 41 minutes adds substantially to the total and pulls the mean above the median. The median remains at the middle of the ordered list.

The difference between the mean and median is not itself a complete description of the data. Here, the ordered values show why the mean is larger: most times are between 12 and 18 minutes, and one much longer time extends the pattern toward the right. A histogram or dotplot would let you see the overall shape directly.

The median’s relative stability is also worth noticing. If the 41-minute wait were replaced by 21 minutes, the total would be \(140\) minutes and the mean would be \(140/9\approx15.56\) minutes. The median would still be 15 minutes. Changing a high observation by 20 minutes changed the mean by \(20/9\approx2.22\) minutes, but it did not change the median in this example.

How Much Does One High Value Matter?

A second comparison makes the difference especially clear. Consider observations that represent the minutes spent completing a fictional set of practice tasks. There are eleven students’ times, and the largest time is much higher than the others. We can compare the original list with a version in which that largest time is lower.

Worked Example: Comparing Two Sets of Practice Times

The original ordered times, in minutes, are \(2, 3, 3, 4, 4, 5, 5, 6, 6, 7,\) and \(29\). For a comparison, replace 29 with 9, leaving the other ten times unchanged.

Plan. Calculate the mean and median for both lists. Because each list has eleven observations, the median is the sixth value. Compare how much each measure changes when the largest observation falls by 20 minutes.

Do. In the original list, the first ten values sum to \(45\), so the full total is \(45+29=74\). Its mean is

$$ \bar{x}=\frac{74}{11}\approx 6.73\text{ minutes} $$

The sixth value is 5 minutes, so the median is 5 minutes. In the comparison list, the total is \(45+9=54\), and the mean is

$$ \bar{x}=\frac{54}{11}\approx 4.91\text{ minutes} $$

The sixth value is still 5 minutes, so the comparison list’s median is also 5 minutes. The mean decreased by \(6.73-4.91=1.82\) minutes (rounded), while the median did not change.

Conclude. The largest value has a strong effect on the mean because the mean incorporates its full numerical size. The median stays at 5 minutes because the sixth position remains unchanged. In the original list, the high value of 29 minutes helps pull the mean above the median.

This example shows why the mean is called nonresistant to extreme values: it can be pulled by a value far from the rest of the data. The median is called resistant because it is usually less affected by a few unusually high or low observations. “Resistant” does not mean unchangeable. If enough observations change, or if the middle positions shift, the median can change too.

When describing a distribution with the SOCS framework from The SOCS Framework for Describing Distributions, use the graph to discuss shape and report a suitable measure of center. For a right-skewed distribution, the median is often a useful description of a typical observation because it is less influenced by the long tail. The mean is still informative when the arithmetic average is meaningful for the question; the two measures simply answer somewhat different questions.

Compare the Pattern Across Shapes

The mean–median relationship often follows a helpful pattern. In a roughly symmetric distribution, the two measures are often close. In a right-skewed distribution, the mean is often greater than the median. In a left-skewed distribution, the mean is often less than the median because a few relatively small values can pull the average downward.

These are typical relationships, not definitions of the shapes. You should not label a distribution right-skewed solely because its mean is larger than its median. The graph provides direct evidence about shape; the numerical summaries help explain or support the description. In particular, a small difference between the mean and median does not prove symmetry.

Worked Example: Three Shapes, Three Center Comparisons

Imagine three small, fictional data sets of temperatures in degrees Celsius. Each list is ordered. The first has a high-end tail, the second is balanced around its center, and the third has a low-end tail.

PatternOrdered observationsMeanMedian
Right tail1, 2, 3, 4, 1043
Symmetric2, 4, 6, 8, 1066
Left tail1, 7, 8, 9, 1078

Plan. Each set has five observations, so its median is the third value. For each mean, add the five observations and divide by five. Then compare the direction of the difference with the tail in the list.

Do. For the right-tail list, the sum is \(1+2+3+4+10=20\), so the mean is \(20/5=4\), while the median is 3. For the symmetric list, the sum is \(2+4+6+8+10=30\), so the mean is \(30/5=6\), equal to the median of 6. For the left-tail list, the sum is \(1+7+8+9+10=35\), so the mean is \(35/5=7\), while the median is 8.

Conclude. In the first list, the large value at the right pulls the mean above the median. In the balanced second list, they are equal. In the third list, the small value at the left pulls the mean below the median. These calculations illustrate a common relationship between tail direction and the two measures of center.

Small data sets are useful for seeing the arithmetic, but their patterns may not represent what happens in every larger data set. With only a few observations, one value can have a large influence, and the shape may be difficult to judge. Use the display and the context, not just a comparison of two numbers, to describe a distribution.

Using the Relationship Carefully

A good explanation links the numerical comparison to the observed values. Instead of writing only “the mean is greater than the median,” identify the group and variable, report the measures with units, and explain that the larger values in the right tail raise the average. If a graph is available, describe the tail that supports that explanation.

For example, after the service-desk calculation, a careful statement is: “For the nine students, the mean wait was about 17.78 minutes and the median was 15 minutes. The distribution has a tail toward longer waits, and the 41-minute wait pulls the mean above the median.” This connects the measures to the context and gives evidence for the explanation.

The mean–median relationship is a clue, not a substitute for looking at the data. A distribution with several peaks or an unusual mixture of groups may have a mean and median that do not tell a simple story about its shape. Likewise, a high mean could reflect values that are all fairly spread out rather than a clearly right-skewed pattern. Refer to the histogram, dotplot, or other suitable display when you make a claim about skew.

As you learned in What a Boxplot Cannot Show, a boxplot summarizes selected locations and spread but does not reveal every feature of a distribution’s shape. If you need to judge the direction or length of a tail, a histogram or dotplot can show that pattern more directly. The mean and median add numerical information; they do not replace the graph.

Common Mistakes and AP Exam Tips

  • Reversing the tail direction. “Right-skewed” means the tail extends toward larger values, not that most observations are on the right. State the direction of the tail explicitly.
  • Claiming every right-skewed distribution must have mean greater than median. This is a common pattern, not a guaranteed rule. Use cautious wording such as “the mean is often greater” and check the actual graph and values.
  • Calling the mean the most common value. The mean is an arithmetic average, not necessarily an observed value or the mode. Explain what the calculation represents.
  • Assuming the median is unaffected by any change. A single high value may leave the middle position unchanged, but changes to enough observations can move the median. Say that the median is generally more resistant, not that it never changes.
  • Describing shape using only two summaries. A mean above a median can support a right-skew explanation, but a graph is better evidence for the distribution’s shape. Do not claim a tail direction without considering the display.
  • Leaving out context or units. For full-credit communication, name the group and variable, report the mean and median in the correct units, and connect the difference to the tail or high values shown.

A complete explanation usually has three parts: report the mean and median in context, describe the relevant tail using the graph or ordered values, and explain that values in that tail affect the mean more strongly than the median. Avoid claiming that the relationship proves a cause or applies to every data set.

Key takeaway: A right tail contains relatively large values, and those values can pull the mean above the median. The median is usually more resistant to a few extreme observations. Use both numerical summaries and a graph to describe shape and center carefully.

Check Your Understanding

Use the calculations and the direction of the tail to explain each comparison.

  1. A fictional ordered data set is \(3, 4, 4, 5, 6, 28\). Find its mean and median. Which is larger, and how does the high value affect the comparison?
  2. Why can replacing one very large observation with a smaller value change the mean while leaving the median unchanged?
  3. A distribution is approximately symmetric, and its mean is close to its median. What does that relationship suggest, and why is it not proof of symmetry?
  4. For a left-skewed distribution, which measure of center is often smaller: the mean or the median? Explain how the tail can affect the mean.
  5. What additional evidence should you examine before claiming that a distribution is right-skewed because its mean exceeds its median?