Welcome to Analysing Research!
Ever wondered how psychologists make sense of the heaps of information they collect during an experiment or an observation? Welcome to the tool kit of psychological investigation! In this chapter, you will learn how to sort data, calculate summary statistics, check if your findings are reliable and valid, and spot hidden biases that might ruin an experiment. Don't worry if maths or scientific terms seem a bit daunting at first — we will break down every single concept step by step.
1. Types of Data: The Building Blocks
When psychologists conduct research, the information they gather can be classified in two main ways: by its form (numbers vs. words) and by its source (first-hand vs. second-hand).
Quantitative vs. Qualitative Data
Quantitative Data: Numerical data that can be counted or measured in numbers.
Examples: The number of words recalled on a memory test, reaction time in seconds, or ratings on a scale of 1 to 10.
• Strengths: It is straightforward to summarise, put into graphs, and analyse using statistics. It allows for quick, objective comparisons between different groups.
• Weaknesses: It lacks depth and context. It tells us what happened, but fails to explain why participants felt or acted the way they did.
Qualitative Data: Non-numerical, descriptive data expressed in words, descriptions, or meanings.
Examples: Word-for-word transcripts of open-ended interviews, diary entries, or written descriptions of emotions.
• Strengths: It provides rich, detailed insight into complex human experiences, emotions, and personal viewpoints.
• Weaknesses: It is difficult to summarise and compare statistically. It is also more open to subjective interpretation and bias from the researcher.
Memory Trick: Remember that Quanti-tative has an 'N' for Numbers, while Quali-tative has an 'L' for Letters / Language!
Primary vs. Secondary Data
Primary Data: First-hand data collected directly by the researcher specifically for their own investigation.
Example: A psychologist designs an experiment on sleep and gathers memory scores from their own participants.
• Strength: The data fits the exact aim and hypothesis of the study, and the researcher controls the quality and procedure.
• Weakness: It can be time-consuming and expensive to plan, recruit participants, and collect.
Secondary Data: Second-hand data that already exists, having been collected previously by someone else for another purpose.
Examples: Government census figures, police crime statistics, or published results from previous studies.
• Strength: It is quick and cheap to obtain, and provides access to huge historical or national datasets that would be impossible to collect alone.
• Weakness: The data might not perfectly match the researcher's specific research question, and the researcher cannot verify how accurately it was gathered.
Key Takeaway: Quantitative = numbers, Qualitative = words. Primary = gathered yourself, Secondary = gathered by someone else.
2. Descriptive Statistics: Making Sense of the Numbers
Once raw quantitative data is gathered, psychologists use descriptive statistics to summarise patterns.
Measures of Central Tendency (Averages)
Measures of central tendency identify the "typical" or central value in a dataset.
1. The Mean: The arithmetic average.
• How to calculate: Add up all the scores in the dataset and divide by the total number of scores: \(\text{Mean} = \frac{\Sigma x}{N}\).
• Advantage: It is the most sensitive measure because it includes every single piece of data in the calculation.
• Disadvantage: It can be heavily distorted by extreme scores (anomalies or outliers).
2. The Median: The middle score when all data values are put in numerical order from lowest to highest.
• How to calculate: Line up all numbers from lowest to highest. If there is an odd number of scores, choose the middle one. If there is an even number of scores, find the mean of the two middle numbers.
• Advantage: It is not distorted by extreme outliers.
• Disadvantage: It does not take into account the exact value of every score in the dataset.
3. The Mode (and Modal Class): The most frequently occurring score in a dataset. In grouped data, the category with the highest frequency is called the modal class.
• Advantage: It is very easy to find and is the only average that can be used for categorical (nominal) data (e.g., favorite colour).
• Disadvantage: A dataset might not have a mode, or it might have more than one mode (bimodal), which makes it less useful as a summary.
Measures of Dispersion (Spread)
The Range: The difference between the highest and lowest scores in a dataset.
• How to calculate: \(\text{Range} = \text{Highest value} - \text{Lowest value}\) (or \(\text{Highest} - \text{Lowest} + 1\)).
• Advantage: Extremely simple and quick to calculate.
• Disadvantage: It only considers the two extreme values and completely ignores the distribution of the numbers in between. A single outlier can make the spread seem misleadingly wide.
Visualising Data: Graphs and Charts
In the exam, you may need to choose, draw, or interpret different graphs. Keep these essential rules in mind:
• Frequency Tables: Simple grids displaying categories alongside raw counts (frequencies).
• Bar Charts: Used for discrete or categorical data (e.g., comparing score averages across experimental conditions). Crucial exam rule: You must leave gaps between the bars to show the categories are separate!
• Histograms: Used for continuous data divided into equal numeric intervals. Unlike bar charts, the bars must touch each other.
• Scatter Diagrams (Scatter plots): Used to show a relationship or correlation between two co-variables (one variable on the \(x\)-axis, one on the \(y\)-axis). Individual data points are plotted as dots/crosses. Do not join the dots with a line!
• Pie Charts & Percentages: Circular charts showing proportions. To calculate a percentage from raw data: \(\text{Percentage} = \frac{\text{part}}{\text{whole}} \times 100\).
Key Takeaway: Mean uses all numbers, Median is the middle number, Mode is the most common. Bar charts have gaps; histograms do not!
3. Reliability: Is the Study Consistent?
Reliability refers to the consistency and replicability of a measuring tool or study. If you repeat the investigation under the exact same conditions, does it produce the same results?
Types of Reliability and How to Test Them
1. Internal Reliability: Consistency between different parts within the same test or assessment tool.
• How to assess it: The split-half method. A test is split into two halves (such as odd-numbered vs. even-numbered questions). If participants get similar scores on both halves, the test has high internal reliability.
2. External Reliability: Consistency of a test or measure across different occasions over time.
• How to assess it: The test-retest method. The same test is given to the same group of participants on two separate occasions. If the two sets of scores are similar, the test has high external reliability.
3. Inter-Rater (Inter-Observer) Reliability: The level of agreement between two or more independent observers who are watching the same event or behaviour.
• How to assess it: Two observers watch the same behaviour independently using clearly defined, standardised behavioral categories. Afterward, their recorded scores are correlated. A strong positive correlation shows high inter-rater reliability.
• Exam Tip: Simply having two observers does not automatically guarantee reliability. They must record data independently and compare their scores to check for strong agreement!
Key Takeaway: Reliability = Consistency. Think: "If I test it again, will I get the same answer?"
4. Validity: Is the Study Accurate?
Validity refers to the accuracy of a research study or measuring tool — does it genuinely measure what it claims to measure, and do the findings reflect reality?
Types of Validity to Know
1. Construct Validity: The extent to which a test or measurement tool truly captures the psychological concept (the "construct") it is supposed to measure.
Example: If a researcher claims that counting how many times someone taps their desk measures "anxiety", this has low construct validity because tapping could just mean boredom.
2. Ecological Validity: The extent to which research findings can be generalized from the experimental setting to real-life, everyday situations.
Example: A memory experiment conducted in an artificial, silent laboratory with lists of random syllables may have low ecological validity because it does not reflect how memory works in daily life.
3. Population Validity: The extent to which findings from a research sample can be generalized to the wider target population.
Example: If a study on stress only tests 18-year-old male university students, it has low population validity because the results might not apply to people of different ages or genders.
Common Mistake Alert: Never confuse Reliability with Validity! A broken bathroom scale that always weighs you \(5\text{ kg}\) too light is reliable (it is consistent every day), but it is not valid (it is inaccurate).
5. Sources of Bias and Extraneous Variables
Bias occurs when systematic errors or expectations distort research findings. Psychologists work hard to identify and eliminate these biases.
Key Sources of Bias
1. Researcher / Experimenter Bias: When a researcher's personal expectations, beliefs, or presence consciously or unconsciously influence the design, execution, or interpretation of a study.
2. Observer Bias: When an observer's prior knowledge, expectations, or stereotypes lead them to notice or record only certain behaviours during an observation.
• How to reduce it: Use clearly defined (operationalised) behavioural checklists, train observers, use "blind" observers (who don't know the hypothesis), and calculate inter-rater reliability.
3. Demand Characteristics: Subtle clues or cues in an experiment that give away the true aim of the study, causing participants to change their natural behaviour to match (or sabotage) what they think the researcher wants.
• How to reduce it: Use a single-blind design (where participants are unaware of which experimental condition they are in or the true aim) or use mild deception where ethically acceptable.
4. Social Desirability Bias: The tendency for participants to give untruthful or exaggerated answers on questionnaires and interviews to make themselves look good or avoid embarrassment.
• How to reduce it: Keep all questionnaires completely anonymous and assure participants of full confidentiality.
5. Gender Bias: When research exaggerates differences between genders, favours one gender, or inappropriately assumes findings from one gender apply equally to all genders.
• How to reduce it: Use representative, gender-balanced samples and avoid making unsupported generalisations about all people.
6. Cultural Bias (including Ethnocentrism): Judging behaviour based on the standards and values of one's own culture, or assuming that findings from one cultural group apply universally across all cultures worldwide.
• How to reduce it: Conduct research using cross-cultural samples rather than relying solely on participants from a single cultural background.
Key Takeaway: Bias threatens the validity of research. Using blind procedures, standardised checklists, anonymity, and representative sampling helps keep studies fair and accurate.
6. Top Exam Tips & Pitfalls to Avoid
• Context is King: When answering exam questions based on a given scenario, never give generic textbook definitions alone. Always apply your answer using the exact names, variables, behaviours, or numbers provided in the case study.
• Median with Even Datasets: Remember to always arrange the raw data in order from lowest to highest first! If you have an even number of values, take the mean of the two middle values.
• Drawing Graphs: Always include an informative title and label both axes clearly (including units of measurement). Leave gaps between the bars for bar charts, but touch the bars together for histograms!
• Scatter Plots: Plot the data using clear dots or crosses — do not join the dots with a continuous line unless specifically asked to add a line of best fit.