Unit 5.2: Correlation
Welcome to the heart of Unit 5! In the previous chapter, we looked at scatterplots to visually describe relationships between two quantitative variables. Now, it is time to put a specific number to that relationship. This number is called the correlation coefficient, or simply \(r\).
Think of correlation as the "strength and direction meter" for a straight-line relationship. It helps us move past just saying a relationship looks "strongish" and gives us a precise statistical value to work with.
What is the Correlation Coefficient \(r\)?
The correlation coefficient \(r\) is a numerical measure that describes the direction and strength of the linear relationship between two quantitative variables.
Important Note: For the AP exam starting in 2027, you are not required to calculate \(r\) by hand using a complex formula. You will use technology (your graphing calculator) or be provided the value in a computer output table. Your job is to understand what that number actually tells you!
The Essential Properties of \(r\)
To master this chapter, you need to know these "rules of the road" for correlation:
- The Range: The value of \(r\) is always between \(-1\) and \(1\), inclusive. In math terms: \(-1 \leq r \leq 1\).
- The Sign (Direction):
- If \(r > 0\), there is a positive association (as \(x\) increases, \(y\) tends to increase).
- If \(r < 0\), there is a negative association (as \(x\) increases, \(y\) tends to decrease).
- The Magnitude (Strength):
- If \(r\) is close to \(1\) or \(-1\), the points lie very close to a straight line. This indicates a strong linear relationship.
- If \(r\) is close to \(0\), the points are very spread out. This indicates a weak linear relationship.
- If \(r = 1\) or \(r = -1\), the points fall perfectly on a straight line.
Analogy: Think of \(r\) like a magnet. The closer the number is to the "poles" (\(1\) or \(-1\)), the more strongly the data points are pulled toward a single straight line. If the number is at \(0\), the "magnet" isn't working, and the points are just drifting aimlessly.
Crucial Limitations: What \(r\) Can and Cannot Do
Don't worry if this feels like a lot of rules—most students find these much easier to remember once they see them in action. Here are the "fine print" details you must know for the exam:
1. It only measures LINEAR relationships.
If your scatterplot looks like a curve (a "U" shape or a rainbow), the correlation coefficient \(r\) is not an appropriate measure. You could have a very clear curved relationship, but \(r\) might still be \(0\). Always look at your scatterplot first!
2. It has no units.
Because \(r\) is calculated using standardized values, it doesn't matter if you measure height in inches or centimeters; the correlation \(r\) will stay exactly the same.
3. Switching \(x\) and \(y\) doesn't change \(r\).
If you find the correlation between "study hours" (\(x\)) and "test scores" (\(y\)), and then you swap them, the value of \(r\) remains identical. Correlation doesn't care which variable comes first.
4. It is NOT resistant to outliers.
Just one single point that is far away from the linear pattern can drastically increase or decrease your \(r\) value. One outlier can make a strong relationship look weak, or a weak relationship look strong.
Quick Review: Interpreting Strength
While there isn't a "universal" rule for what counts as strong, here is a general guide for AP Statistics:
- \(|r| \approx 0.0\) to \(0.3\): Weak
- \(|r| \approx 0.4\) to \(0.7\): Moderate
- \(|r| \approx 0.8\) to \(1.0\): Strong
The Golden Rule: Correlation Does Not Establish Causation
This is perhaps the most important sentence in all of statistics! Just because two variables have a high correlation (\(r\) close to \(1\) or \(-1\)), it does not mean that one variable causes the change in the other.
Example: There is a high positive correlation between ice cream sales and the number of shark attacks. Does eating ice cream cause shark attacks? Of course not! A "lurking variable"—hot weather—causes both to increase.
Key Takeaway: Only a well-designed experiment with random assignment can establish causation. Correlation only shows that a relationship exists.
Common Pitfalls to Avoid
Mistake 1: Confusing Correlation with Slope.
Students often think a steep line means a high correlation. Not true! You can have a very flat line (small slope) where all the points are perfectly on the line (\(r = 1\)), or a very steep line where the points are scattered everywhere (\(r = 0.2\)). Correlation is about closeness to the line, not the steepness of the line.
Mistake 2: Using "Correlation" for Categorical Data.
You can only calculate \(r\) for two quantitative variables (numbers like height, weight, or temperature). You cannot find the "correlation" between eye color and gender.
Summary Checklist
Before moving on to Section 5.3: Linear Regression Models, make sure you can:
- Identify the direction (positive/negative) from the sign of \(r\).
- Identify the strength based on how close \(r\) is to \(1\) or \(-1\).
- Explain why a specific \(r\) value might be misleading (e.g., if the data is curved or has outliers).
- Explain that a high correlation does not mean \(x\) causes \(y\).
Pro Tip: On a Free Response Question (FRQ), if you are asked to describe a relationship, always mention four things: Direction, Form (Linear), Strength, and Unusual Features (Outliers). Correlation \(r\) helps you quantify two of those: Direction and Strength!