Welcome to Comparing Data Sets & Drawing Conclusions!
Have you ever tried to decide which of two football players is performing better, or which class did better in a recent maths test? To answer questions like these fairly, you cannot just look at one single number. In statistics, comparing two or more sets of data is a vital skill that helps us make smart decisions, spot patterns, and understand the world around us.
Don't worry if this seems tricky at first! By the end of these notes, you will have a simple, reliable toolkit to compare any two data sets with complete confidence.
1. The Golden Rule of Comparing Data Sets
Whenever an exam question asks you to "compare two distributions" or "compare and comment on two sets of data", you must ALWAYS write about at least TWO key features:
• 1. A Measure of Average (Central Tendency): This tells us what is typical or average about the data (e.g., Mean or Median).
• 2. A Measure of Spread (Dispersion): This tells us how spread out or consistent the data is (e.g., Range, Interquartile Range (IQR), or Standard Deviation).
Memory Trick (The "A-S-C" Rule):
• A - Average (Who scored higher / did better on average?)
• S - Spread (Who was more consistent / less varied?)
• C - Context (Always include the real-life meaning and units!)
Quick Review: Consistency
A smaller measure of spread (smaller Range, smaller IQR, or smaller Standard Deviation) means the data values are closer together. In statistical terms, we say the data is more consistent or less variable.
Key Takeaway: Never compare just the averages! A complete statistical comparison always includes one average and one measure of spread, written in the context of the problem.
2. Choosing the Right Statistical Measures
Which average and spread should you choose? It depends on the shape of your data and whether any extreme values (outliers) are present.
Scenario A: Symmetrical Data (No Extreme Outliers)
• Best Average: Mean (\( \bar{x} \))
• Best Spread: Standard Deviation (\( \sigma \)) or Range
Why? The mean uses every single piece of data, making it very powerful when the data is evenly balanced.
Scenario B: Skewed Data or Data with Outliers
• Best Average: Median (\( Q_2 \))
• Best Spread: Interquartile Range (\( \text{IQR} = Q_3 - Q_1 \))
Why? Extreme values (very high or very low numbers) distort the mean and range. The median and IQR are resistant to outliers because they only focus on the middle values.
Did you know? If a billionaire walks into a cafe with ten students, the mean wealth of the room suddenly jumps into millions of pounds! However, the median wealth barely changes at all. That is why the median is preferred when dealing with skewed income data.
Key Takeaway: Use Mean & Standard Deviation for symmetrical data with no outliers. Use Median & IQR when the data is skewed or contains outliers.
3. How to Structure Your Comparison in Exams
To gain full marks on CCEA GCSE Statistics questions, use this simple 3-step formula for every statement:
Step 1: State the values clearly. Give the numerical figures for both groups.
Step 2: Use comparative words. Words like higher, lower, greater, less, more consistent, or more spread out.
Step 3: State the real-life context. Mention the units, people, or items being measured.
Worked Example: Comparing Test Scores
Suppose Class A has a Median = \(68\%\) and an \(\text{IQR} = 12\%\).
Class B has a Median = \(54\%\) and an \(\text{IQR} = 26\%\).
• Comparing the Average:
"Class A had a higher median score (\(68\%\)) than Class B (\(54\%\)), which means Class A performed better on average."
• Comparing the Spread:
"Class A had a smaller interquartile range (\(12\%\)) than Class B (\(26\%\)), which means test scores in Class A were more consistent and less spread out."
Key Takeaway: Always quote both values and clearly state what that means in the real-world context of the question.
4. Comparing Different Types of Statistical Diagrams
A. Comparing Box Plots (Box-and-Whisker Diagrams)
Box plots are fantastic for direct comparisons when drawn on the same scale.
• Look at the Middle Line: This is the median (\( Q_2 \)). The box plot with its median line further to the right represents higher values on average.
• Look at the Width of the Box: This represents the \(\text{IQR} = Q_3 - Q_1\). A narrower box indicates a more consistent middle \(50\%\) of data.
• Look at Total Width (Whiskers): This shows the overall \(\text{Range} = \text{Maximum} - \text{Minimum}\).
• Look at Skewness:
- If the median is in the middle of the box, the data is symmetrical.
- If the median is closer to the lower quartile (\(Q_1\)), the data is positively skewed.
- If the median is closer to the upper quartile (\(Q_3\)), the data is negatively skewed.
B. Comparing Cumulative Frequency Graphs
• Median (\(50\%\) position): Read across from \(50\%\) of the total frequency to each curve and down to the horizontal axis. A curve that sits further to the right generally has higher values.
• Steepness: A steeper curve in the middle indicates that many values are clustered together, showing a smaller IQR (more consistent data).
C. Comparing Back-to-Back Stem-and-Leaf Diagrams
• Back-to-back stem-and-leaf diagrams share a central stem.
• Watch out: For the left-hand group, leaves read from right to left (inside out)!
• Find the median and range for both sides to draw your conclusions.
Key Takeaway: Visual comparisons allow you to easily estimate medians, spreads, and general shapes before calculating exact figures.
5. Understanding Shapes and Skewness
When discussing distributions, their shape tells an important story:
• Symmetrical Distribution: The distribution forms a bell-like curve. The mean, median, and mode are approximately equal (\( \text{Mean} \approx \text{Median} \approx \text{Mode} \)).
• Positively Skewed (Right-tailed): The bulk of the data is clustered at the lower end, with a long tail stretching to the higher values. In this case: \( \text{Mode} < \text{Median} < \text{Mean} \).
• Negatively Skewed (Left-tailed): The bulk of the data is clustered at the higher end, with a long tail stretching towards the lower values. In this case: \( \text{Mean} < \text{Median} < \text{Mode} \).
Analogy: Imagine your toes! Look at your left foot from above: the big toe is on the right and the little toes tail off to the left (negatively skewed). On your right foot, the big toe is on the left and the little toes tail off to the right (positively skewed).
Key Takeaway: Skewness indicates whether data bunches at the low end (positive skew) or high end (negative skew).
6. Common Mistakes to Avoid
• Mistake 1: Merely listing numbers without comparing them.
Wrong: "Group A has a median of 20. Group B has a median of 15."
Correct: "Group A has a higher median (\(20\)) than Group B (\(15\)), meaning Group A had greater times on average."
• Mistake 2: Forgetting the real-world context.
Always mention what the numbers represent (e.g., marks, minutes, heights, kilograms).
• Mistake 3: Confusing smaller spread with "worse" performance.
A smaller range or IQR simply means the results are more consistent or closer together; it does not automatically mean higher or lower performance.
• Mistake 4: Claiming that a higher average means every single member did better.
An average describes the group as a whole. Even if Group A has a higher median than Group B, some individuals in Group B might still have scored higher than individuals in Group A.
7. Final Summary Checklist
When answering a comparison question, check off these points:
✔ Have I compared one average (Mean or Median)?
✔ Have I compared one measure of spread (Range, IQR, or Standard Deviation)?
✔ Have I used comparative words (higher, lower, more consistent)?
✔ Have I stated actual numerical values from the data?
✔ Have I written my conclusion clearly in context?