Welcome to the Normal Distribution!
Have you ever noticed that most adult shoe sizes cluster around the middle (like UK size \(7\) to \(9\)), while very small sizes (like size \(3\)) and very large sizes (like size \(14\)) are quite rare? Or that most students score near the class average on an exam, with only a few getting exceptionally low or high marks?
In statistics, this classic pattern is called the Normal Distribution (often called the bell-shaped curve). It is one of the most powerful and widely used tools in statistics because so many natural measurements follow this exact shape. Don't worry if this sounds intimidating at first—by the end of these notes, you will know exactly how it works, how to read it, and how to use it to solve exam questions with confidence!
1. What is the Normal Distribution?
The Normal Distribution is a continuous probability distribution that is perfectly symmetrical. It models data that clusters around a central average value, with values becoming less frequent the further you move away from the centre.
Key Properties of the Normal Distribution
Every normal distribution shares these essential features:
• Bell-Shaped Curve: It rises smoothly to a single peak in the middle and slopes downwards on both sides.
• Symmetrical: If you draw a vertical line straight down the middle, the left half is a mirror image of the right half.
• Mean = Median = Mode: In a perfect normal distribution, the average (mean), middle value (median), and most frequent value (mode) are all exactly equal and sit right at the centre peak.
• Total Area Under the Curve is \(1\) (or \(100\%\)): The entire area underneath the curve represents all possible outcomes, which equals \(1\) (or \(100\%\) of the data).
• Asymptotic Tails: The curve spreads out infinitely to the left and right, getting closer and closer to the horizontal axis, but it never actually touches it.
• Defined by Two Parameters: Every normal distribution is completely described by two numbers:
1. The mean (\(\mu\)) – determines the centre position of the curve.
2. The standard deviation (\(\sigma\)) – determines the spread or width of the curve.
Did You Know?
The normal distribution was historically called the "Gaussian distribution", named after the German mathematician Carl Friedrich Gauss, who used it to analyse astronomical data in the early 1800s!
Key Takeaway:
A normal distribution is a symmetrical, bell-shaped curve where the mean, median, and mode are equal, and the total area under the curve is always equal to \(1\) (or \(100\%\)).
2. The \(68 - 95 - 99.7\%\) Rule (The Empirical Rule)
One of the most important tools you will need for CCEA GCSE Statistics is the Empirical Rule. This rule tells us approximately what percentage of data falls within a certain number of standard deviations (\(\sigma\)) away from the mean (\(\mu\)).
The Three Key Percentages to Memorise:
• Approximately \(68\%\) of all data lies within \(1\) standard deviation of the mean:
Between \(\mu - 1\sigma\) and \(\mu + 1\sigma\)
• Approximately \(95\%\) of all data lies within \(2\) standard deviations of the mean:
Between \(\mu - 2\sigma\) and \(\mu + 2\sigma\)
• Approximately \(99.7\%\) of all data lies within \(3\) standard deviations of the mean:
Between \(\mu - 3\sigma\) and \(\mu + 3\sigma\)
Breaking Down the Symmetrical Slices
Because the curve is symmetrical, we can split these regions into easy-to-use slices on either side of the mean (\(\mu\)):
• Since \(68\%\) lies within \(1\sigma\), exactly half of that, \(34\%\), lies between \(\mu\) and \(\mu + 1\sigma\), and \(34\%\) lies between \(\mu - 1\sigma\) and \(\mu\).
• Between \(1\sigma\) and \(2\sigma\) from the mean lies approximately \(13.5\%\) on each side (calculated as \(\frac{95\% - 68\%}{2} = 13.5\%\)).
• Outside \(2\sigma\) from the mean (the outer tails) lies only \(5\%\) in total, meaning roughly \(2.5\%\) in each extreme tail!
Memory Aid: The "Rule of Thumb"
Think of standard deviations like ripples in a pond around a stone (the mean):
• 1st ripple (\(\pm 1\sigma\)): catches the main crowd (\(68\%\))
• 2nd ripple (\(\pm 2\sigma\)): catches almost everyone (\(95\%\))
• 3rd ripple (\(\pm 3\sigma\)): catches practically everything (\(99.7\%\))
Worked Example: Applying the Rule
Question: The heights of a large group of adult females are normally distributed with a mean of \(\mu = 165\text{ cm}\) and a standard deviation of \(\sigma = 6\text{ cm}\).
(a) Find the range of heights that contains the middle \(68\%\) of women.
(b) What percentage of women are taller than \(177\text{ cm}\)?
Step-by-step Solution:
Part (a):
The middle \(68\%\) lies within \(1\) standard deviation of the mean (\(\mu \pm 1\sigma\)).
Lower boundary = \(\mu - 1\sigma = 165 - 6 = 159\text{ cm}\)
Upper boundary = \(\mu + 1\sigma = 165 + 6 = 171\text{ cm}\)
Answer: The middle \(68\%\) of women are between \(159\text{ cm}\) and \(171\text{ cm}\).
Part (b):
First, check how many standard deviations \(177\text{ cm}\) is above the mean:
\(\mu + 2\sigma = 165 + 2(6) = 165 + 12 = 177\text{ cm}\).
We know that \(95\%\) of women have heights between \(\mu - 2\sigma\) (\(153\text{ cm}\)) and \(\mu + 2\sigma\) (\(177\text{ cm}\)).
The remaining percentage outside this range is \(100\% - 95\% = 5\%\).
Because the curve is symmetrical, half of this \(5\%\) is below \(153\text{ cm}\) and half is above \(177\text{ cm}\):
Percentage above \(177\text{ cm} = \frac{5\%}{2} = 2.5\%\).
Answer: Approximately \(2.5\%\) of women are taller than \(177\text{ cm}\).
Key Takeaway:
Remember the three magic numbers: \(68\%\) (\(\pm 1\sigma\)), \(95\%\) (\(\pm 2\sigma\)), and \(99.7\%\) (\(\pm 3\sigma\)). Use the symmetry of the curve to find values in the tails.
3. Standardised Scores (\(z\)-Scores)
Imagine you scored \(75\) in a History exam and \(80\) in a Maths exam. Did you perform better in Maths? Not necessarily! If the Maths exam was very easy (high class average) and the History exam was very hard (low class average), your History score might actually be more impressive.
To compare values from different normal distributions fairly, we convert them into a standardised score (also known as a \(z\)-score).
The Standardisation Formula
A standardised score tells us exactly how many standard deviations a value lies above or below the mean.
\(z = \frac{x - \mu}{\sigma}\)
Where:
• \(x\) = the raw value you are testing
• \(\mu\) = the mean of the distribution
• \(\sigma\) = the standard deviation of the distribution
How to Interpret a \(z\)-Score
• Positive \(z\)-score (\(z > 0\)): The value is above the mean.
• \(z = 0\): The value is exactly equal to the mean.
• Negative \(z\)-score (\(z < 0\)): The value is below the mean.
• Large magnitude (e.g., \(z > +3\) or \(z < -3\)): The value is very unusual or an outlier.
Worked Example: Comparing Test Results
Question:
Sarah scored \(68\) on a Physics test where \(\mu = 60\) and \(\sigma = 4\).
She scored \(74\) on a Chemistry test where \(\mu = 65\) and \(\sigma = 6\).
In which subject did Sarah perform relatively better?
Step-by-step Solution:
Step 1: Calculate the \(z\)-score for Physics:
\(z_{\text{Physics}} = \frac{x - \mu}{\sigma} = \frac{68 - 60}{4} = \frac{8}{4} = +2.0\)
(Sarah's Physics score is \(2\) standard deviations above the average.)
Step 2: Calculate the \(z\)-score for Chemistry:
\(z_{\text{Chemistry}} = \frac{x - \mu}{\sigma} = \frac{74 - 65}{6} = \frac{9}{6} = +1.5\)
(Sarah's Chemistry score is \(1.5\) standard deviations above the average.)
Step 3: Compare:
Since \(+2.0 > +1.5\), Sarah's standardised score is higher in Physics.
Answer: Relative to the rest of the class, Sarah performed better in Physics.
Key Takeaway:
Standardised scores (\(z = \frac{x - \mu}{\sigma}\)) allow fair comparisons between different data sets by measuring performance relative to the mean and standard deviation.
4. Comparing Different Normal Curves
Changing the parameters \(\mu\) and \(\sigma\) alters how the normal curve looks on a graph:
Effect of Changing the Mean (\(\mu\))
• If two distributions have the same standard deviation but different means, they have the exact same shape and width, but one is shifted horizontally along the axis to the left or right.
Effect of Changing the Standard Deviation (\(\sigma\))
• If two distributions have the same mean but different standard deviations, they are centred at the same place, but:
– A smaller \(\sigma\) produces a tall, narrow peak (data is tightly clustered around the mean).
– A larger \(\sigma\) produces a flatter, wider curve (data is more spread out).
• Remember: the total area under both curves remains exactly equal to \(1\)!
Key Takeaway:
The mean (\(\mu\)) controls the location/centre of the curve, while the standard deviation (\(\sigma\)) controls the spread/height of the curve.
5. Real-World Applications & Limitations
When is the Normal Distribution a Good Model?
The normal distribution is commonly used to model:
• Physical characteristics of living things (human height, arm span, animal weights).
• Manufacturing processes (mass of cereal in boxes, volume of drink in bottles, length of bolts).
• Standardised test scores and IQ scores.
• Measurement errors in scientific experiments.
Limitations to Keep in Mind:
• Real data is rarely perfectly normal: Real data may have slight skewness or contain extreme outliers.
• Negative values in theoretical models: A theoretical normal curve extends from \(-\infty\) to \(+\infty\), but real variables like height or weight cannot be negative.
6. Common Mistakes to Avoid in Exams
• Forgetting negative signs for \(z\)-scores: If a score is below the mean, \(x - \mu\) is negative, so \(z\) must be negative!
• Confusing the mean and standard deviation: Always double-check which number is \(\mu\) and which is \(\sigma\) before substituting into the formula.
• Assuming all symmetrical distributions are normal: A distribution can be symmetrical without having the specific bell curve percentages (\(68-95-99.7\%\)).
• Forgetting that total area = \(100\%\): Always use the symmetry of the curve (halving the remaining percentages) when calculating tail proportions.
Quick Review Checklist
Can you answer these key revision questions? If so, you are ready for this topic!
✔ What three measures of average are equal in a normal distribution? (Mean, Median, Mode)
✔ What percentage of data lies within \(1\) standard deviation of the mean? (\(\approx 68\%\))
✔ What percentage of data lies within \(2\) standard deviations of the mean? (\(\approx 95\%\))
✔ What percentage of data lies within \(3\) standard deviations of the mean? (\(\approx 99.7\%\))
✔ What does a \(z\)-score of \(-1.5\) tell you? (The value is \(1.5\) standard deviations below the mean.)
✔ What is the total area underneath a normal distribution curve? (Exactly \(1\) or \(100\%\))