Introduction to Evaluating Research
In Psychology, we don't just accept research findings at face value. We have to be "psychological detectives." When you read about a study—like Milgram’s work on obedience or Raine’s brain scans—you need to ask: "Is this study actually good quality?"
To answer this, we use a specific set of tools called evaluation criteria. These help us decide how much we can trust the results and how we can apply them to the real world. In your Pearson Edexcel International AS Level exams, you will often be asked to evaluate studies using these concepts. Let’s break them down into easy-to-remember parts!
1. Generalisability: Can we apply this to others?
Generalisability refers to the extent to which the findings of a study can be applied to the wider target population or to different settings. If a study has high generalisability, it means the results aren't just true for the people in the experiment, but are likely true for most people.
What makes generalisability high or low?
- Sample Size: Larger samples are usually more generalisable because they are more likely to represent the whole population.
- Sample Diversity: If a study only uses male students (like Milgram (1963) initially did), it might have low generalisability to females or older people.
- Sampling Method: Random or stratified sampling usually leads to better generalisability than opportunity sampling.
Key Takeaway: Always look at who was in the study. If the group was too specific (e.g., only 6-year-old twins in Brendgen et al. (2005)), ask yourself if the results would be the same for adults or single children.
2. Reliability: Is it consistent?
Reliability is all about consistency. If we repeated the study exactly the same way tomorrow, would we get the same results? Think of a bathroom scale: if you step on it and it says \(70kg\), then step off and back on and it says \(75kg\), the scale is unreliable.
How do psychologists ensure reliability?
- Standardised Procedures: This means every participant has the exact same experience. For example, in Milgram's study, the "prods" used by the experimenter (e.g., "The experiment requires that you continue") were scripted so every participant heard the same thing.
- Controls: Keeping the environment the same (e.g., the same room, the same equipment) helps ensure that the results aren't just a "one-off" fluke.
Quick Tip: If a study is "replicable" (easy to copy), it is more likely to be checked for reliability.
3. Validity: Is it truthful?
Validity asks: "Are we actually measuring what we claim to be measuring?" A study can be reliable (consistent) but still be invalid (not measuring the right thing).
There are three main types you need to know for your exam:
A. Internal Validity
This is about control. Did the Independent Variable (IV) really cause the change in the Dependent Variable (DV), or was it something else (an extraneous variable)? If a study has a lot of "noise" or distractions, its internal validity is low.
B. Predictive Validity
Does the result of the study or test accurately predict what will happen in the future? For example, if a test for "aggression" has high predictive validity, people who score highly on it should actually behave aggressively in real life later on.
C. Ecological Validity
This is about the setting and the task. Does the study represent real life?
Example: Bartlett (1932) asked people to remember a story called "The War of the Ghosts." Some argue this has low ecological validity because we don't usually sit in labs memorising strange folk tales in our daily lives.
Don't worry if this seems tricky! Just remember: Internal = Control; Ecological = Realism.
4. Objectivity and Subjectivity: Facts vs. Feelings
These two terms describe how much the researcher’s own opinions influenced the results.
Objectivity
Objectivity is when research is "unbiased." The findings are based on hard facts and data that anyone can see.
How to be objective: Use quantitative data (numbers), like the voltage levels in Milgram’s study or the brain activity in Raine et al. (1997). A brain scan doesn't have an "opinion"—it just shows what it shows!
Subjectivity
Subjectivity is when the findings are influenced by personal feelings, interpretations, or prejudice.
When does this happen? Usually when collecting qualitative data. For example, in thematic analysis, one researcher might think a participant's interview answer belongs in one category ("theme"), while another researcher might disagree. This relies on human judgment.
Key Takeaway: High objectivity makes a study more scientific, but subjectivity can sometimes provide deeper, more detailed "rich" data.
5. Credibility: Is the whole thing believable?
Credibility is like the "reputation" of the research. A study is considered credible if it is believable and convincing. Credibility is usually high if the study is objective, valid, and reliable. If a study has very low ecological validity (it's too artificial) or uses a very small, biased sample, its overall credibility might be questioned.
Did you know? Using triangulation (using different methods to check the same thing) can make research more credible!
Quick Summary Table
Use this "Cheat Sheet" to help you remember the definitions for your Unit 1 and Unit 2 exams:
| Term | Simple Definition |
| Generalisability | Can the results be applied to people outside the study? |
| Reliability | Are the results consistent? Can the study be repeated? |
| Internal Validity | Did the IV definitely cause the DV? (Was it well-controlled?) |
| Ecological Validity | Is the study realistic/representative of real-world behavior? |
| Objectivity | The research is free from researcher bias (fact-based). |
| Subjectivity | The research involves personal interpretation (feeling-based). |
Common Mistake to Avoid: Don't confuse Reliability with Validity! A clock that is always 5 minutes fast is reliable (it’s consistent) but it is not valid (it doesn’t tell the true time).
Final Check: How to use this in an exam
When you see an 8-mark or 12-mark question asking you to "evaluate" a study (like Moscovici (1969) or Schmolck et al. (2002)), use these terms as your paragraph headings. For example:
"One strength of Schmolck et al. (2002) is its high objectivity. This is because they used scores from standardised memory tests, which provides quantitative data that is not subject to researcher interpretation."