Welcome to Research Methods: Psychology in Context (Paper 2, Section C)

Welcome to one of the most important chapters in your entire AQA A Level Psychology course! Research methods makes up Section C of Paper 2 (Psychology in Context), carrying a massive 48 marks (half of the paper) and worth 33.3% of your total A Level. Furthermore, at least 10% of the overall marks across all three papers test your mathematical and data-handling skills.

Don't worry if maths or scientific terminology feels intimidating at first. We will break every single concept down into clear, bitesize chunks with everyday analogies, clear rules, and practical examples so you can score full marks with confidence.

---

Section 1: Scientific Processes & Experimental Methods

1. Experimental Methods

An experiment is the only research method that allows researchers to establish a direct cause-and-effect relationship by manipulating an Independent Variable (IV) to see its effect on a Dependent Variable (DV).

There are four distinct experimental types:

1. Laboratory Experiments: Conducted in a tightly controlled, artificial environment where the researcher directly manipulates the IV.
Strengths: High control over extraneous variables; high internal validity; easy to replicate.
Weaknesses: Artificial setting leads to low ecological validity and low mundane realism (tasks do not reflect everyday life); risk of demand characteristics.

2. Field Experiments: Conducted in a natural, real-world setting, but the researcher still deliberately manipulates the IV.
Strengths: Higher mundane realism and higher ecological validity than lab studies.
Weaknesses: Loss of control over extraneous variables; harder to replicate; potential ethical issues (e.g., lack of informed consent if participants do not know they are in a study).

3. Natural Experiments: The researcher does not manipulate the IV; the change in the IV occurs naturally (e.g., studying mental health before and after a natural disaster).
Strengths: Allows research on unique, real-life situations that would be unethical to engineer artificially.
Weaknesses: Naturally occurring events happen rarely; participants cannot be randomly allocated to conditions.

4. Quasi-Experiments: The IV is based on an existing, pre-set individual difference or characteristic (e.g., age, gender, neurodivergence). It is not manipulated by anyone.
Strengths: Allows comparisons between distinct types of people under controlled conditions.
Weaknesses: Participants cannot be randomly allocated, meaning confounding participant variables cannot be eliminated.

Common Examiner Trap: Do not confuse natural and quasi-experiments! In a natural experiment, the IV is an environmental event that happened on its own. In a quasi-experiment, the IV is a fixed participant trait (like age or gender).

2. Aims and Hypotheses

Aim: A general statement explaining what the researcher intends to investigate.
Null Hypothesis: Predicts that there will be no significant difference or no significant relationship between variables (any observed change is purely due to chance).
Alternative / Experimental Hypothesis: Predicts that there will be a significant difference or relationship.

There are two types of experimental hypotheses:

Directional Hypothesis (One-tailed): States the specific direction of the outcome (e.g., "Participants who drink coffee will recall significantly more words than participants who drink water").
When to use: ONLY when previous published research or theory suggests a particular direction.

Non-directional Hypothesis (Two-tailed): States that there will be a difference, but does not state which direction (e.g., "There will be a significant difference in the number of words recalled between participants who drink coffee and participants who drink water").
When to use: When there is NO prior research, or when past research findings are contradictory.

Top Tip for Full Marks — Operationalisation: Hypotheses must be fully operationalised (clearly defined and objectively measurable). Never just say "caffeine affects memory." Say: "Drinking 200 ml of caffeinated coffee compared to 200 ml of water will result in a higher score on a 20-word recall test."

3. Experimental Designs

How do we arrange our participants across our experimental conditions?

1. Independent Groups: Different participants take part in each experimental condition.
Strength: No order effects (participants cannot get tired, bored, or better through practice).
Weakness: Participant variables (individual differences like IQ or age) can confound the results.
Control: Random allocation (e.g., drawing names out of a hat) to distribute individual differences evenly.

2. Repeated Measures: The same participants take part in all conditions.
Strength: Eliminates participant variables; requires fewer total participants.
Weakness: Order effects (fatigue, boredom, or practice effects) and higher chance of demand characteristics.
Control: Counterbalancing using an \(AB\) or \(BA\) design (half the participants do Condition \(A\) then \(B\); the other half do Condition \(B\) then \(A\)).

3. Matched Pairs: Different participants are used in each condition, but they are pre-tested and matched in pairs on key characteristics (e.g., identical IQ, age, or reading ability). One member goes to Condition \(A\), the other to Condition \(B\).
Strength: Reduces participant variables without causing order effects.
Weakness: Time-consuming and difficult to match people perfectly.

4. Variables and Controls

Extraneous Variables (EVs): Any nuisance variables that might affect the DV if not controlled (e.g., room lighting, background noise).
Confounding Variables (CVs): Variables that change systematically alongside the IV, meaning we cannot know whether the IV or the CV caused the change in the DV.
Demand Characteristics: Clues in the study that reveal the aim to participants, leading them to alter their natural behaviour (controlled using a single-blind design, where participants do not know which condition they are in).
Investigator Effects: Any conscious or unconscious cues from the researcher that influence participant behaviour (controlled using standardisation and a double-blind design, where neither the participant nor the researcher running the trial knows the condition/hypothesis).
Standardisation: Keeping all instructions, environments, and procedures identical for every participant.
Randomisation: Using chance (e.g., shuffling word lists, flipping coins) to eliminate researcher bias in the design of materials.

5. Sampling Techniques

The target population is the entire group the researcher is interested in studying. The sample is the smaller group selected to take part.

Random Sampling: Every member of the target population has an equal chance of selection (e.g., all names placed into a random number generator). Free from researcher bias, but can be unrepresentative by chance.
Systematic Sampling: Selecting every \(n\)th person from a sampling frame (e.g., every 5th name on an alphabetical list). Objective and avoids researcher bias.
Stratified Sampling: The sample reflects the exact proportions of identified sub-groups (strata) in the target population. Highly representative, but time-consuming.
Opportunity Sampling: Selecting anyone who is available and willing at the time of the study. Quick and economical, but highly unrepresentative and biased toward a specific location/time.
Volunteer (Self-Selected) Sampling: Participants choose to take part by responding to adverts or posters. Easy to gather, but suffers from volunteer bias (attracts a specific, highly motivated type of person).

6. Pilot Studies

A pilot study is a small-scale practice run of an investigation using a small number of participants. Its purpose is to check instructions, test timings, trial measuring instruments, and fix design flaws before investing significant time and money into the full investigation.

Key Takeaway: Good experimental research isolates cause-and-effect by operationalising variables, standardising procedures, counterbalancing order effects, and choosing a representative sample.

---

Section 2: Non-Experimental Methods & Observational Design

1. Observational Techniques & Design

Observations involve watching and recording spontaneous behaviour. They can be set up along three main spectrums:

Naturalistic vs Controlled: Observed in an unaltered, everyday setting vs observed in a structured laboratory setting.
Covert vs Overt: Participants are unaware they are being watched (secret) vs aware they are being watched (open).
Participant vs Non-participant: The researcher joins the group being studied vs remains an outside observer.

Observational Design Features:

Behavioural Categories: Breaking down a continuous target behaviour into clear, observable, measurable, and non-overlapping components (e.g., operationalising "aggression" into kicking, pushing, shouting).
Event Sampling: Counting the number of times a specific target behaviour occurs every time it happens.
Time Sampling: Recording target behaviours at predetermined time intervals (e.g., noting what the participant is doing every 30 seconds).

2. Self-Report Techniques: Questionnaires & Interviews

Questionnaires:
Closed questions: Fixed response options (e.g., rating scales, yes/no). Produce quantitative data (easy to analyse statistically, but lacks depth).
Open questions: Allow participants to answer freely in their own words. Produce qualitative data (rich in detail, but difficult to statistically analyse and compare).

Interviews:
Structured: Set, pre-determined questions read aloud identically to every participant. Standardised and easy to replicate, but inflexible.
Unstructured: Conversational with no set questions; questions are developed based on participant responses. Rich and flexible, but difficult to analyse and prone to interviewer bias.
Semi-structured: A set list of core questions with the freedom to ask follow-up questions for clarification.

3. Correlations

A correlation measures the strength and direction of an association between two continuous co-variables (e.g., hours of revision and exam scores).

Positive Correlation: As one co-variable increases, the other co-variable increases.
Negative Correlation: As one co-variable increases, the other co-variable decreases.
Zero Correlation: No relationship exists between the co-variables.

Crucial Distinction: Experiments manipulate an IV to measure an effect on a DV, allowing us to establish causality (cause-and-effect). Correlations simply measure the relationship between two co-variables without manipulation. Correlation does NOT equal causation because an unmeasured third variable (an intervening variable) might be driving the relationship.

4. Case Studies & Content Analysis

Case Studies: Detailed, in-depth investigations of a single individual, small group, institution, or event. They use idiographic triangulation (combining interviews, observations, and psychological tests). They produce rich, detailed qualitative data, but have low population validity and cannot be generalised.
Content Analysis: Indirect observation of human communication through media (books, adverts, diaries, speeches).
- Coding: Turning qualitative data into quantitative counts by tallying target categories.
- Thematic Analysis: A qualitative process that identifies recurring themes or ideas across the material.

Key Takeaway: Non-experimental methods give us rich qualitative insights (case studies) and help us spot patterns (correlations), but cannot directly prove cause-and-effect.

---

Section 3: Reliability, Validity, Ethics & Scientific Principles

1. Ethics: The BPS Code of Conduct

Psychological research in the UK must adhere to the British Psychological Society (BPS) Code of Ethics and Conduct:

Informed Consent: Participants must understand the true aims and procedures before agreeing. If full consent is impossible initially, researchers can use presumptive consent (asking a similar group), prior general consent, or retrospective consent (gained during debriefing).
Deception: Deliberately withholding information or misleading participants must be kept to an absolute minimum, justified, and followed by a full debrief.
Protection from Harm: Participants must not experience physical or psychological harm greater than they would in everyday life.
Confidentiality & Anonymity: Personal data must be protected; participants must remain unidentifiable (e.g., using numbers or initials).
Right to Withdraw: Participants have the right to leave the study at any time and can demand that their data be destroyed post-study.
Debriefing: After the study, participants must be given a full explanation of the true aims, offered psychological support, and allowed to ask questions.

2. Reliability (Consistency)

Reliability is how consistent a measuring tool or procedure is. If repeated under identical conditions, does it produce identical results?

Ways to Assess Reliability:
1. Test-Retest Reliability: The same test is administered to the same participants on two separate occasions. If the two sets of scores produce a correlation coefficient of \(+0.80\) or higher, reliability is confirmed.
2. Inter-Observer / Inter-Rater Reliability: Two or more independent observers record behaviour simultaneously using identical behavioural categories. Their recorded scores are correlated; a correlation coefficient of \(+0.80\) or higher indicates strong inter-observer reliability.

Improving Reliability: Standardise all instructions; operationalise behavioural categories clearly to avoid overlap; use closed questions instead of ambiguous open questions.

3. Validity (Accuracy & Truthfulness)

Validity refers to whether a test actually measures what it claims to measure, and whether the findings represent real life.

Internal Validity: Whether the observed changes in the DV are solely caused by the IV (free from confounding variables and demand characteristics).
External Validity: The extent to which findings can be generalised beyond the study setting:
- Ecological validity: Generalisability to everyday real-world settings.
- Temporal validity: Generalisability across different time periods/eras.
- Population validity: Generalisability to other groups of people.

Ways to Assess Validity:
1. Face Validity: An expert inspects the measuring tool "on the face of it" to see if it looks like it measures what it claims to.
2. Concurrent Validity: Comparing scores on a new measuring tool against an established, validated test taken by the same participants. A correlation coefficient of \(+0.80\) or higher confirms concurrent validity.

4. Features of Science

To be considered a true science, psychological research must meet core criteria:

Objectivity & Empirical Method: Gathering direct, observable evidence that is completely free from personal researcher bias or expectations.
Replicability: Procedures must be standardised and recorded in detail so other researchers can repeat them and verify findings.
Falsifiability (Karl Popper): A theory is not scientific unless it can theoretically be proven false through empirical testing.
Theory Construction & Hypothesis Testing: Formulating broad principles that explain observed behaviours, from which clear, testable predictions (hypotheses) can be deduced and tested.
Paradigms & Paradigm Shifts (Thomas Kuhn): A paradigm is a universally accepted set of assumptions and methods shared by a scientific discipline. A paradigm shift occurs when significant conflicting evidence (anomalies) overthrows the old paradigm in a scientific revolution (e.g., the shift from behaviourism to cognitive psychology).

5. Peer Review & the Economy

Peer Review: The independent, anonymous assessment of scientific research papers by other expert psychologists before publication. Its purpose is to:
1. Allocate research funding objectively.
2. Validate the quality, originality, and methodology of research.
3. Prevent fraudulent or flawed research from entering scientific journals.

Psychology and the Economy: Psychological research helps the wider economy by:
• Developing effective treatments for mental disorders (e.g., SSRIs, CBT for depression) which reduces NHS costs and reduces workplace absenteeism.
• Research into the role of the father and attachment showing that flexible shared parental leave supports maternal return to work and boosts household economic productivity.

6. Reporting Psychological Investigations

Scientific papers follow a standard format:

1. Title: Clear and concise.
2. Abstract: A 150–200 word summary covering aims, hypotheses, method, results, and conclusions.
3. Introduction: Background literature review, rationale, aims, and hypotheses.
4. Method: Detailed enough for exact replication (Design, Sample, Materials/Apparatus, Procedure, Ethics).
5. Results: Descriptive statistics (tables, graphs, measures of central tendency) and inferential statistical test results.
6. Discussion: Interpretation of results, real-world applications, limitations, and future research directions.
7. References: Standard academic citations (e.g., Flanagan, C. (2020). Psychology Revision Guide. London: Oxford Press.).

Key Takeaway: Science relies on public scrutiny (peer review), consistent tools (reliability \(\ge +0.80\)), accurate measures (validity), and strict moral codes (BPS ethics).

---

Section 4: Data Handling, Descriptive & Inferential Statistics

1. Types of Data & Levels of Measurement

Quantitative: Numerical data.
Qualitative: Non-numerical, descriptive data in words.
Primary Data: Collected first-hand by the researcher specifically for the current investigation.
Secondary Data: Information gathered by someone else prior to the investigation (e.g., government statistics, journal articles).
Meta-analysis: A statistical technique that pools and combines the findings of multiple secondary studies investigating the same topic to calculate an overall effect size.

Levels of Measurement (The "NOIR" Scale):

1. Nominal Data: Data placed into named, separate, discrete categories or frequency counts (e.g., tally of pass/fail, types of pets owned: dog, cat, bird).
2. Ordinal Data: Data that is ranked or placed in order, but the intervals between units are not equal or standardised (e.g., 1st, 2nd, 3rd place in a race; rating happiness on a scale of 1–10).
3. Interval Data: Data measured on a continuous numerical scale with standardised, equal intervals between units (e.g., temperature in degrees Celsius, time in seconds, standardized test scores).

2. Descriptive Statistics: Central Tendency & Dispersion

Measures of Central Tendency (Averages):
Mean: The mathematical average (sum of all scores divided by total number of scores). Most sensitive measure because it includes every piece of data, but easily distorted by extreme anomalies (outliers).
Median: The middle score when all data points are arranged in numerical order. Unaffected by extreme outliers, but ignores the value of individual scores.
Mode: The most frequently occurring score. The only average suitable for nominal category data, but can be uninformative if there are multiple modes.

Measures of Dispersion (Spread of Scores):
Range: The difference between the highest and lowest score (conventionally calculated as \(\text{Highest} - \text{Lowest} + 1\)). Easy to calculate, but heavily distorted by one extreme score.
Standard Deviation (SD): Measures the average distance/spread of all scores away from the mean. A low standard deviation shows that scores are clustered tightly around the mean; a high standard deviation shows that scores are widely spread out.

3. Distributions: Normal and Skewed Curves

Normal Distribution: A symmetrical, bell-shaped curve where the Mean, Median, and Mode all sit exactly at the same central peak. In a normal distribution, \(68.26\%\) of scores fall within \(\pm 1\text{ SD}\) of the mean, and \(95.44\%\) fall within \(\pm 2\text{ SD}\).

Positive Skew: The distribution tail stretches out to the right (scores cluster toward the lower end, e.g., an extremely hard exam where most people score low).
Order on the graph from left to right: \(\text{Mode} < \text{Median} < \text{Mean}\).

Negative Skew: The distribution tail stretches out to the left (scores cluster toward the higher end, e.g., an extremely easy test where most people score high).
Order on the graph from left to right: \(\text{Mean} < \text{Median} < \text{Mode}\).

4. Inferential Testing: Choosing the Correct Statistical Test

In the exam, when asked to justify a statistical test, you must state all 3 criteria:

1. Are you testing for a difference or a correlation/association?
2. What is the experimental design (unrelated / independent groups vs related / repeated measures / matched pairs)?
3. What is the level of measurement (Nominal, Ordinal, or Interval)?

The AQA 8-Test Decision Grid:

Nominal Data:
• Unrelated Design \(\implies\) Chi-square (\(\chi^2\))
• Related Design \(\implies\) Sign Test
• Test of Correlation/Association \(\implies\) Chi-square (\(\chi^2\))

Ordinal Data:
• Unrelated Design \(\implies\) Mann-Whitney \(U\)
• Related Design \(\implies\) Wilcoxon \(T\)
• Test of Correlation \(\implies\) Spearman's Rho (\(r_s\))

Interval Data (Parametric tests):
• Unrelated Design \(\implies\) Unrelated \(t\)-test
• Related Design \(\implies\) Related \(t\)-test
• Test of Correlation \(\implies\) Pearson's \(r\)

Criteria for Using a Parametric Test:

To use a parametric test (Unrelated \(t\)-test, Related \(t\)-test, or Pearson's \(r\)), the data must satisfy three conditions:
1. Data must be at the Interval (or ratio) level.
2. Data must be drawn from a population with a Normal distribution.
3. Homogeneity of variance (the spread of scores must be roughly equal in both conditions).

5. How to Calculate the Sign Test (Step-by-Step)

The Sign Test is the only statistical test you may be asked to calculate by hand in your exam. Follow these 4 easy steps:

Step 1: For each participant, subtract Condition \(B\) score from Condition \(A\) score. Record whether the change is positive (\(+\)) or negative (\(-\)).
Step 2: Count the total number of \(+\) signs and \(-\) signs. If any participant showed zero change (a tie), exclude them and subtract them from your total \(N\).
Step 3: Find your calculated value (\(S\)). \(S\) is simply the frequency of the less common sign.
Step 4: Compare your calculated \(S\) value to the critical value from the statistical table for your given significance level (usually \(p \le 0.05\)) and your adjusted \(N\).
• For the Sign test: The result is statistically significant if Calculated \(S \le \text{Critical value}\).

6. Probability, Significance & The "Rule of R"

• In psychology, the standard level of significance is \(p \le 0.05\) (meaning there is a \(5\%\) or less probability that the results occurred by chance).
• A stricter level of \(p \le 0.01\) (\(1\%\)) is used when the research involves human risk (e.g., testing new medical drugs) or when replicating a past study.

The "Rule of R" Memory Trick for Statistical Tables:
How do you know if your calculated value needs to be higher or lower than the critical value?

• If the test has an 'R' in its name (Spearman's, Pearson's, Chi-square, Unrelated \(t\), Related \(t\)):
\(\implies\) Calculated value \(\ge\) Critical value for significance.
• If the test does NOT have an 'R' in its name (Mann-Whitney, Wilcoxon, Sign test):
\(\implies\) Calculated value \(\le\) Critical value for significance.

7. Statistical Decision Errors: Type I and Type II

Type I Error (\(\alpha\) error / "False Positive"): Rejecting a null hypothesis that is actually true (claiming a significant difference/effect exists when it does not). This happens when your significance level is too lenient (e.g., \(p = 0.10\)).
Type II Error (\(\beta\) error / "False Negative"): Accepting/retaining a null hypothesis that is actually false (failing to notice a real difference/effect). This happens when your significance level is set too strictly (e.g., \(p = 0.001\)).

Quick Memory Aid:
Type I = Optimistic error ("I found something!" ... but you didn't).
Type II = Pessimistic error ("There's nothing here..." ... but there was).

Key Takeaway: Choose your statistical test using the 3 dimensions (Difference/Correlation \(\times\) Design \(\times\) Measurement Level). Use \(p \le 0.05\) as standard, remember the "Rule of R", and double-check for ties when calculating the Sign Test.