Welcome to Geographic Statistics and Data Interpretation

In Geography, data is our evidence. Whether we are looking at how a city grows or how a river flows, we use numbers to prove our theories. This chapter will guide you through the essential statistical tools you need for your IB Geography exams and your Internal Assessment (IA). Don't worry if you aren't a "math person"—these tools are all about finding patterns in the world around us!

Note: For help with mapping and scales, please see the chapter "Location, maps and scale."

1. The Basics: Summary Statistics

Before we can do complex analysis, we need to summarize our data. These calculations help us describe a large set of numbers with just one or two values.

Totals, Averages, and Frequencies

Totals: The sum of all values in a data set. Useful for looking at total rainfall or total population.

Mean (Average): The sum of all values divided by the number of values \( (n) \).
Formula: \( \bar{x} = \frac{\sum x}{n} \)

Median: The middle value when the data is placed in order. It is great for avoiding the "skewing" effect of extreme numbers (anomalies).

Mode: The value that appears most frequently in a data set.

Frequency: How often a specific value occurs. We often use frequency tables to organize raw data from fieldwork.

Ranges and Densities

Range: The difference between the highest (maximum) and lowest (minimum) values.
Formula: \( \text{Range} = \text{Maximum} - \text{Minimum} \)

Density: How much of something exists within a specific area. Geographers use this for population (people per \( km^2 \)) or drainage density in a river basin.
Formula: \( \text{Density} = \frac{\text{Total Number or Amount}}{\text{Total Land Area}} \)

Percentages and Ratios

Percentages: Used to compare parts of a whole.
Formula: \( \text{Percentage} = (\frac{\text{Part}}{\text{Whole}}) \times 100 \)

Ratios: A way to compare the size of two different groups (e.g., a doctor-to-patient ratio of \( 1:500 \)).

Quick Review: Summary statistics tell you "what" the data looks like, but they don't tell you "why" patterns exist.

2. Measures of Correlation: Finding Relationships

Correlation tells us if two variables are linked. For example, does the quality of health decrease as poverty increases?

Spearman’s Rank Correlation Coefficient \( (r_s) \)

This is a favorite in IB Geography. It tests the strength and direction of a relationship between two sets of ranked data.

The Result: The answer will always be between \( +1 \) and \( -1 \).
- \( +1 \): A perfect positive correlation (as one goes up, the other goes up).
- \( -1 \): A perfect negative correlation (as one goes up, the other goes down).
- \( 0 \): No relationship at all.

Chi-squared Test \( (\chi^2) \)

This test is used to see if there is a significant difference between the observed frequencies (what you actually counted) and the expected frequencies (what you would expect if there was no pattern). It is often used to see if the location of something is "random" or influenced by a specific factor.

3. Measures of Concentration and Dispersion

These tools help geographers describe how things are spread out across space.

Nearest Neighbour Index (NNI)

This measures the spatial distribution of points (like settlements or shops). It tells us if they are:
1. Clustered: Grouped together (NNI close to \( 0 \)).
2. Random: No clear pattern (NNI of \( 1.0 \)).
3. Regular/Uniform: Evenly spaced out (NNI up to \( 2.15 \)).

Location Quotients (LQ)

This shows the concentration of a particular characteristic in an area compared to a larger region. For example, is a specific industry more concentrated in one city than in the rest of the country? An LQ higher than \( 1.0 \) means that characteristic is more "concentrated" there than average.

4. Geographic Indices and Ratios

The IB syllabus requires you to understand specific indices used to measure development, inequality, and resource use.

The Human Development Index (HDI)

A composite index (meaning it combines different types of data) used to measure development. It includes:
- Health: Life expectancy at birth.
- Education: Mean years of schooling.
- Living Standards: GNI per capita.

The Gini Coefficient and Lorenz Curve

These measure inequality (usually income).
- The Lorenz Curve is a graph showing the cumulative percentage of income against the cumulative percentage of the population.
- The Gini Coefficient is a number calculated from that graph. \( 0 \) represents perfect equality (everyone has the same income), and \( 1 \) (or \( 100\% \)) represents perfect inequality.

The Dependency Ratio

This shows the relationship between the "dependent" population (those too young or too old to work) and the "economically active" population.
Formula: \( \frac{\text{Population (0–14) + Population (65+)}}{\text{Population (15–64)}} \times 100 \)

Ecological Footprint

A measure of the area of biologically productive land and water required to produce the resources a population consumes and to absorb its waste.

5. Data Interpretation and Evaluation

Calculating the numbers is only half the battle. In Paper 2 and Paper 3, you must interpret what those numbers mean.

Classifying and Analysing

When looking at data, try to:
- Identify Trends: Is the data generally increasing, decreasing, or staying the same over time?
- Identify Anomalies: Look for "outliers"—data points that don't fit the general pattern. Why are they there? Is it a recording error or a special local factor?
- Make Generalizations: Can you sum up the pattern in one sentence? (e.g., "As distance from the CBD increases, building height generally decreases.")

Evaluating Methodology

In your IA and exam questions, you may be asked to evaluate your methods. Consider:
- Sample Size: Was it big enough to be representative?
- Bias: Did you only collect data at one time of day or in one location?
- Accuracy: Were the instruments used (like flow meters or questionnaires) reliable?

Key Takeaways for the Exam:

- Always show your working out in calculations.
- Use specific units (e.g., \( km^2 \), \( \% \), or \( US\$ \)).
- When describing patterns, always quote highest and lowest figures from the data provided to support your answer.