Welcome to the World of Data!

In Psychology, once we have finished an experiment or a questionnaire, we are often left with a mountain of numbers. On their own, these numbers don't tell us much. This is where Descriptive Statistics come in! These are tools that help us summarise, organise, and describe our data so we can see the "big picture" before we move on to more complex math.

In this chapter, we will look at List A of your research methods requirements. These are the fundamental skills you will use across all your practical investigations in Unit 1 and Unit 2.


1. The Basics: Percentages, Ratios, and Fractions

Before we get into "Psychology-specific" math, we need to be comfortable with basic numbers. In your exam, you might be asked to convert raw scores from a study into these formats.

Percentages \((\%)\)

Percentages show us a proportion out of 100. They are great for comparing groups of different sizes. For example, if 15 out of 40 people obeyed in a study, what percentage is that?

The Calculation: \(\frac{\text{Part}}{\text{Whole}} \times 100\)

Example: \(\frac{15}{40} \times 100 = 37.5\%\)

Ratios

Ratios compare one amount to another. If 20 participants are "obedient" and 10 are "disobedient," the ratio is \(20:10\). We always simplify this by dividing both sides by the smallest number. Here, \(2:1\).

Fractions

Fractions represent a part of a whole. If 5 out of 20 people remembered a word list, the fraction is \(\frac{5}{20}\), which simplifies to \(\frac{1}{4}\).

Quick Review: Always read the question carefully to see if it asks you to simplify your answer or round it to a certain number of decimal places!


2. Measures of Central Tendency

A "measure of central tendency" is just a fancy way of saying "the average." It’s a single value that represents the typical score in your data set.

The Mean

This is the arithmetic average. You add all the scores together and divide by the total number of scores (\(n\)).

When to use it: It is the most sensitive measure because it uses every single piece of data. However, it can be easily "thrown off" by one or two very high or very low scores (outliers).

The Median

This is the middle value when all scores are placed in order from lowest to highest. If there is an even number of scores, the median is the average of the two middle numbers.

When to use it: It is great for data with outliers because the "extreme" scores at the ends don't change the middle value much.

The Mode

This is the most frequently occurring score in a data set. A set can have one mode, two modes (bimodal), or no mode at all.

When to use it: It is the only measure you can use for "nominal" data (categories). For example, "What is the most common hair colour in the classroom?"

Memory Tip:
MOde = MOst frequent.
MEdian = MEddle (middle).
Mean = The Mean one (because it takes the most work to calculate!).


3. Measures of Dispersion

While central tendency tells us the average, Dispersion tells us how spread out the scores are. Are everyone’s scores similar, or are they all over the place?

The Range

This is the simplest measure. You take the highest score and subtract the lowest score.

Pros: Very easy to calculate.
Cons: It only looks at the two most extreme scores and ignores everything in the middle.

Standard Deviation (SD)

The Standard Deviation tells us the average amount that each score differs (deviates) from the mean. A low SD means the scores are clustered closely around the mean. A high SD means the scores are widely spread out.

The Formula: In your exam, you will be given this formula in the booklet:

\(s = \sqrt{\frac{\sum(x - \bar{x})^2}{n - 1}}\)

Don't Panic! Let's break it down:
1. \(x\) is each individual score.
2. \(\bar{x}\) (x-bar) is the mean of the scores.
3. \(\sum\) means "the sum of."
4. \(n\) is the number of participants.
5. Basically, you find the difference between each score and the mean, square those differences, add them up, divide by \(n-1\), and then take the square root.

Key Takeaway: Standard Deviation is much more powerful than the range because it uses every single score in the data set to tell us about the spread.


4. Data Tables

Psychologists use tables to keep data organized. You need to know two types:

Frequency Tables

These record how often a certain score or category occurs. Usually, these have "Tally" columns.

Example:

Score: 5 | Tally: /// | Frequency: 3

Summary Tables

These appear in the "Results" section of a report. They don't show every single raw score; instead, they show the calculated descriptive statistics (like the Mean and SD) for different groups.


5. Graphical Presentation

Graphs make data visual. For your exam, you must know when to use a Bar Chart versus a Histogram. This is a common area where students lose marks!

Bar Charts

Used for discrete data (data that fits into separate categories). For example, "Type of animal" or "Condition A vs. Condition B."

  • The bars must not touch.
  • The x-axis (bottom) shows the categories.
  • The y-axis (side) usually shows the frequency or the mean score.

Histograms

Used for continuous data (data on a scale that can be broken down into smaller and smaller parts, like time or height).

  • The bars must touch each other.
  • The x-axis shows "equal intervals" (like 0-9 seconds, 10-19 seconds).

Common Mistake: Forgetting to label your axes! Always label the x and y axes and give your graph a clear, descriptive title.


6. Normal and Skewed Distributions

When we plot data on a graph, the "shape" of the data tells us a story.

Normal Distribution

This is the famous "bell curve." It is perfectly symmetrical. In a perfect normal distribution, the mean, median, and mode are all the same and sit right in the middle.

Skewed Distributions

Sometimes data is "leaning" to one side.

  • Positive Skew: Most scores are low, with a few very high scores pulling the "tail" to the right. (Think of a very hard exam where most people failed but a few geniuses got 100%).
  • Negative Skew: Most scores are high, with a few very low scores pulling the "tail" to the left. (Think of an easy exam where almost everyone got an A).

Did you know? In a skewed distribution, the mean is pulled the furthest toward the "tail," while the mode stays at the highest peak.


Summary Checklist: Are you exam-ready?

1. Can you calculate a percentage and simplify a ratio?
2. Do you know which measure of central tendency to use for nominal data? (Hint: It’s the mode!)
3. Can you explain why Standard Deviation is better than the Range?
4. Do you remember that bars must touch in a histogram but must not touch in a bar chart?
5. Can you identify a "tail" pointing to the right as a positive skew?

Note: For how to use these scores in inferential tests like the Wilcoxon or Chi-Squared, see the chapter on List B.