Introduction: Can We Trust the Results?
Imagine you are using a bathroom scale to check your weight. If you step on it three times and get three completely different numbers, you’d probably throw the scale away! In Psychology, we have the same problem. When we conduct a study, we need to know if our results are consistent and if they actually measure what they claim to measure. This is where reliability and validity come in.
In this chapter, we will look at the "big labels" we use to judge how good a piece of research is. These are the tools you will use to evaluate every study you encounter in your International A Level course.
1. Reliability: The Power of Consistency
Reliability refers to how consistent a measurement is. If a finding is "reliable," it means that if we did the study again in the exact same way, we would get the same (or very similar) results.
Think of it like a recipe. If a recipe is reliable, every time you follow the instructions, the cake should taste exactly the same. In Psychology, we achieve this through standardisation—keeping every procedure, instruction, and environment identical for all participants.
Quick Review: How to check for reliability
Scientists look for a strong positive correlation (usually \(r \ge +0.80\)) when they repeat a test. If the results are the same twice, the test is reliable.
2. Validity: The Search for Truth
Validity is about accuracy. It asks: "Are we actually measuring what we think we are measuring?" A study can be reliable (consistent) but totally invalid (inaccurate).
Analogy: Imagine a clock that is exactly 10 minutes fast. It is reliable because it is always 10 minutes fast (it’s consistent). However, it is not valid because it isn't telling you the actual time.
The syllabus requires you to know three specific types of validity:
A. Internal Validity
This asks if the Independent Variable (IV) is the only thing affecting the Dependent Variable (DV). If a study has high internal validity, we are confident that the "cause and effect" we found is real and not caused by "messy" factors like extraneous variables.
B. Predictive Validity
This is the degree to which a test can predict future performance or behavior. For example, if a test for "aggression" has high predictive validity, then people who score highly on it should actually get into more fights in the future.
C. Ecological Validity
This is about the setting. Does the behavior we see in a laboratory represent how people act in the real world? If a study is too artificial (like memorising random lists of words in a dark room), it might have low ecological validity because real-world memory doesn't work that way.
3. Generalisability: From the Few to the Many
Generalisability is the extent to which we can apply the findings of a study to other people and other situations.
If a study only used 20-year-old male students from London, can we say the results apply to 70-year-old women in Tokyo? If the answer is "no," the study has low generalisability. To have high generalisability, we usually need a representative sample (a group that looks like the wider population).
4. Objectivity vs. Subjectivity
Psychology aims to be a science, which means it values objectivity.
- Objectivity: Being unbiased. Results are based on facts and measurements that anyone can see. For example, counting how many times a person smiles is more objective than "guessing" if they are happy.
- Subjectivity: Being influenced by personal feelings or interpretations. If a researcher "interprets" a dream, their own opinions might change the result. This is often seen as a weakness in scientific research.
5. Credibility: Is it Believable?
Credibility is the overall "believability" of the research. It is like the "reputation" of the study. A study is considered credible if it is:
1. Methodologically sound (good reliability and validity).
2. Peer-reviewed by other scientists.
3. Objective and scientifically rigorous.
6. Common Methodological Issues
Even the best-planned studies can run into trouble. Here are two major issues you need to look out for:
Researcher Effects (and Experimenter Effects)
This happens when the researcher accidentally influences the participants. This could be through their body language, the tone of their voice, or even their gender. For example, participants might answer questions about "prejudice" differently if the researcher is from a minority group.
Social Desirability
Don't worry if this seems like a fancy term; it's basically the "Goody-Two-Shoes" effect. Participants often want to look like "good people," so they might lie on questionnaires or change their behavior to seem more moral, kind, or successful than they actually are. This lowers the validity of the data.
Note: This is often linked to demand characteristics, where participants try to guess the aim of the study and change their behavior to "help" or "hinder" the researcher.
7. Key Takeaways Summary
Reliability = Consistency. If you do it again, do you get the same result?
Validity = Accuracy. Are you measuring what you intended to?
Generalisability = Application. Does this apply to the "real world" and other people?
Objectivity = Facts. Keeping personal bias out of the results.
Social Desirability = Lying to look good. A major threat to the truth of the data.
Quick Tip for the Exam: When you are asked to evaluate a study, always try to use these "GRAVE" pillars: Generalisability, Reliability, Application, Validity, and Ethics!