Welcome to Comparing Distributions!

In the previous chapters, you learned how to describe the distribution of a single quantitative variable. That’s a great start, but in the real world, we rarely look at one group in isolation. We want to know: Do students who sleep 8 hours perform better than those who sleep 5? Is Brand A’s battery life more consistent than Brand B’s?

In this chapter, you will learn how to take those descriptive skills and use them to highlight the similarities and differences between two or more groups. This is one of the most important skills for the AP Statistics exam, especially for the Free-Response Questions (FRQs).

1. Visualizing Multiple Groups

Before we can calculate anything, we need to see the data. When we compare groups, we use specific types of graphs that put the distributions "side-by-side" so our eyes can easily spot differences. Common tools include:

  • Side-by-Side (Parallel) Boxplots: These are the gold standard for comparing centers and variability. They allow you to see the medians and Interquartile Ranges (IQRs) of several groups at once on the same scale.
  • Back-to-Back Stemplots: Perfect for comparing two small datasets. One group’s leaves go to the right, and the other group’s leaves go to the left, sharing the same stem in the middle.
  • Parallel Dotplots: Similar to boxplots, but they show every individual data point. These are great for seeing gaps or clusters.

Quick Review: For a refresher on how to build these individual graphs, see the chapter "Graphical representations for one quantitative variable".

2. The "Big Four" Comparison Framework

When the AP exam asks you to compare distributions, they are looking for a feature-by-feature breakdown. Don't just list the stats for Group A and then list them for Group B. You must use comparative language (words like greater than, less than, or about the same as).

Use the mnemonic S.O.C.V. (or "S.O.C.S.") to make sure you cover everything:

Shape

Are the distributions skewed left, skewed right, or symmetric?
Example: "The distribution of scores for Group A is skewed to the right, while the distribution for Group B is roughly symmetric."

Outliers (Unusual Features)

Does one group have outliers while the other doesn't? Are there major gaps in one distribution?
Example: "Group A has two high outliers at \( 95 \) and \( 98 \), whereas Group B has no outliers."

Center

Compare the typical values. Usually, we compare medians (especially if the data is skewed) or means (\( \bar{x} \)).
Critical Tip: You must use a comparative word! "The median of Group A (\( 75 \)) is higher than the median of Group B (\( 62 \))."

Variability (Spread)

Which group is more "spread out"? You can compare the Range, the IQR (\( Q_3 - Q_1 \)), or the Standard Deviation (\( s \)).
Example: "The scores in Group B show more variability, with an IQR of \( 20 \) points compared to Group A's IQR of only \( 12 \) points."

Key Takeaway: Always include context (units like "inches," "seconds," or "test points") and always use comparative words.

3. Pro-Tips for Success

Don't just "list" – Compare!
If you write "Group A has a median of 10. Group B has a median of 20," you will likely lose points. You must say "The median of Group B is greater than the median of Group A."

Use the same scale
When looking at two different histograms or dotplots, always check the x-axis. If the scales are different, a "wide" looking graph might actually be narrower than it appears!

The "About" Rule
Statistical data is rarely perfect. Use phrases like "roughly symmetric" or "approximately" when describing shapes or estimating values from a graph. It shows you understand that real-world data has "noise."

4. Step-by-Step Example

Scenario: A scientist compares the growth (in cm) of plants using two different fertilizers.

Data Summary:
Fertilizer X: Median = \( 12 \), IQR = \( 4 \), Shape = Skewed Right.
Fertilizer Y: Median = \( 15 \), IQR = \( 5 \), Shape = Symmetric.

How to write the comparison:
"The distribution of plant growth for Fertilizer Y is roughly symmetric, while Fertilizer X is skewed to the right. Fertilizer Y produced a higher median growth (\( 15 \) cm) compared to Fertilizer X (\( 12 \) cm). However, Fertilizer Y also showed slightly more variability in growth, with an IQR of \( 5 \) cm compared to Fertilizer X’s IQR of \( 4 \) cm. Neither distribution showed any outliers."

Did you know?
Comparing distributions is the first step in Inference. While we are just describing differences now, in later units (Unit 3 and 4), you will learn how to calculate if these differences are "statistically significant" or if they just happened by chance!

5. Common Mistakes to Avoid

  • Ignoring Context: If the problem is about heights, use the word "height" and "inches" in your answer.
  • Mixing up Mean and Median: If a distribution is heavily skewed, the median is usually the better measure of center to use for comparison.
  • Forgetting Unusual Features: Always check for outliers using the \( 1.5 \times IQR \) rule if you have the data. If you are looking at a graph and see a dot far away from the rest, mention it!

Quick Review Box:
1. Use S.O.C.V. (Shape, Outliers, Center, Variability).
2. Use comparative language (higher, lower, more variable).
3. Always include context and units.
4. Use side-by-side boxplots or back-to-back stemplots for visual comparisons.