Welcome to Validity and Reliability!
In Psychology, we are always asking two big questions about our research: "Can I trust these results to stay the same?" and "Do these results actually show the truth?" These questions lead us to the two pillars of psychological research: Reliability and Validity. Think of them as the "Quality Control" of science. Don't worry if these terms seem similar at first—by the end of these notes, you'll be able to spot the difference easily!
1. Reliability: The Quest for Consistency
Reliability refers to how consistent or stable a measure is. If a test is reliable, it should produce the same results every time it is used under the same conditions.
The Bathroom Scale Analogy: Imagine you step on a scale and it says you weigh \( 70kg \). You step off and immediately step back on, and it says \( 70kg \). The scale is reliable because it is consistent. If it said \( 70kg \), then \( 65kg \), then \( 75kg \), it would be unreliable!
Types of Reliability you need to know:
A. Replicability
This is the ability for a study to be repeated exactly. To make a study replicable, researchers use standardisation (keeping everything the same for every participant). If the procedure is clear and controlled, other scientists can "replicate" the study to see if they get the same findings. (Note: For more on how we control studies, see the chapter on "Control of variables and standardisation").
B. Test-retest Reliability
This involves giving the same participants the same test at two different times. If the scores are similar both times, the test has high test-retest reliability. For example, if you take a personality test today and again in two weeks, your results should be roughly the same.
C. Inter-rater and Inter-observer Reliability
Sometimes, research involves observers watching behavior. Inter-observer reliability is the extent to which two or more observers agree.
How it works:
1. Two observers watch the same behavior at the same time.
2. They record their data independently.
3. They compare their results. If their data is very similar (usually a correlation of \( +0.80 \) or higher), the study has high reliability.
Quick Takeaway: Reliability = Consistency. If you see the word "consistent" or "repeatable," think Reliability!
2. Validity: The Quest for Truth
Validity is about accuracy. It asks: "Is the researcher actually measuring what they claim to be measuring?" A study can be consistent (reliable) but still be wrong (invalid).
The Dartboard Analogy:
- If you hit the same spot on the edge of the board every time, you are reliable (consistent) but not valid (you missed the bullseye).
- If you hit the bullseye every time, you are both reliable and valid!
Key Concepts in Validity:
A. Ecological Validity
This is how well the findings of a study represent real life.
- High Ecological Validity: A study in a natural environment, like Piliavin et al. (subway Samaritans), where people are in their normal surroundings.
- Low Ecological Validity: A study in a fake or artificial environment, like a laboratory, where the task might be unusual (like the "Eyes Test" in Baron-Cohen et al.).
B. Subjectivity vs. Objectivity
- Objectivity: Taking an unbiased, factual approach. Using biological scans (like Hölzel et al.) or counting specific behaviors makes a study more valid because it isn't based on opinion.
- Subjectivity: A personal viewpoint that may be biased. If a researcher interprets a participant's feelings based on their own opinion, the validity might decrease.
C. Demand Characteristics
This happens when participants guess the aim of the study and change their behavior to "help" the researcher or look better. This ruins validity because the behavior is no longer natural.
D. Generalisability
This refers to how much the findings can be applied to other people or settings. If a study only uses 10-year-old boys from one school, the results might not generalise to adults or people from different cultures. (Note: This is closely linked to "Sampling of participants").
E. Temporal Validity (A Level Only)
This is whether the findings of a study remain true over time. For example, a study on social behavior from the 1950s might have low temporal validity because society has changed so much since then.
Quick Takeaway: Validity = Accuracy/Truth. If the study is "fake" or "biased," it has low validity.
3. Reliability and Validity in Paper 2 Planning
When you are asked to plan a study in Paper 2, you must explain how you will make your plan reliable and valid.
To increase Reliability:
- Use a standardised procedure (everyone gets the same instructions).
- If using observers, use two observers and check for agreement.
To increase Validity:
- Use a natural setting to improve ecological validity.
- Use a single-blind or double-blind design to reduce demand characteristics (where participants don't know the aim).
4. Comparing the Two (Quick Review)
Did you know? It is possible for a study to be reliable but not valid, but it is almost impossible for a study to be valid if it isn't reliable! Reliability is usually the first step toward a good study.
Common Mistakes to Avoid:
1. Don't swap them! If you say a study is "reliable" because it was in a real-life setting, you are wrong—that is ecological validity.
2. Don't just say "it's valid." Always explain why. Is it valid because it was objective? Is it valid because it avoided demand characteristics?
Summary Table
Concept: Reliability
Key Word: Consistency
How to improve: Standardise everything; use multiple observers.
Concept: Validity
Key Word: Accuracy / Truth
How to improve: Reduce bias; use realistic settings; hide the aim from participants.
Final Tip for Success
Whenever you evaluate a Core Study (like Milgram or Bandura), always ask yourself:
1. "Could I repeat this exactly?" (Reliability/Replicability)
2. "Was the behavior natural or forced?" (Ecological Validity)
3. "Was the data based on facts or opinions?" (Objectivity/Subjectivity)