Understanding Risk, Correlation, and Study Design

Welcome to one of the most important chapters in Topic 1! While you’ve been learning about the biology of the heart and blood vessels, this section focuses on the evidence. How do we actually know that smoking causes heart disease? Why do some people worry more about shark attacks than their diet? In these notes, we will break down how scientists design studies, interpret data, and why "risk" isn't always as simple as it seems.

1. What is Risk?

In biology, risk is defined as the probability that a particular unwanted event (like a disease or death) will happen to a person within a certain period of time.

Actual Risk vs. Perceived Risk
One of the trickiest parts of human biology is that our brains aren't very good at calculating actual risk. Perceived risk is how risky we think an activity is, and it often differs from the actual risk (the statistical reality).

People tend to overestimate risk if the event is:
• Involuntary (you have no control over it).
• Dreaded or spectacular (like a plane crash).
• Unfamiliar or caused by "unnatural" things.

People tend to underestimate risk if the event is:
• Voluntary (like smoking or choosing to eat junk food).
• Familiar (like driving a car).
• Long-term (the "unwanted event" won't happen for many years).

Quick Tip: If an exam question asks why someone continues to smoke despite the health risks, mention that they underestimate the risk because the habit is voluntary and the consequences are in the distant future.

2. Interpreting Data: Mortality and Morbidity

To understand the health of a population, scientists look at two main types of data:
1. Mortality: The number of deaths in a given period.
2. Morbidity: The number of people suffering from a particular disease (illness rates).

When looking at data tables or graphs, always check the units. Is it "deaths per 100,000 people" or a "percentage"? This allows us to compare different populations fairly, even if they are different sizes.

3. Correlation vs. Causation

This is a favorite topic for examiners! It is vital to know the difference between these two terms.

Correlation
A correlation exists when there is a change in one variable that is accompanied by a change in another variable. For example, as the amount of saturated fat in a diet increases, the rate of Cardiovascular Disease (CVD) might also increase. This is a positive correlation.

Causation
Causation is when a change in one variable directly causes the change in another. For a correlation to be causal, there must be a biological mechanism that explains the link.

Example: There is a correlation between ice cream sales and shark attacks. Does ice cream cause shark attacks? No! A third variable (hot weather) causes both. Therefore, the correlation is not causal.

Conflicting Evidence
In Topic 1, you will often see conflicting evidence. One study might say caffeine increases risk, while another says it has no effect. This usually happens because studies use different sample groups or different methods. Science is a process of constant evaluation!

4. Evaluating Study Design

Not all scientific studies are created equal. To get valid and reliable data, a study must be designed carefully.

Sample Selection and Bias
The people chosen for a study (the sample) must represent the whole population.
Random selection is used to avoid bias. If a scientist only chooses fit athletes for a study on heart health, the results won't apply to the general public.
• The sample should represent different ages, genders, and backgrounds.

Sample Size
The larger the sample size, the better!
• A small sample (e.g., 5 people) might show a trend just by chance.
• A large sample (e.g., 10,000 people) reduces the effect of anomalies and makes the results more reliable.

Validity and Reliability
Validity: Does the study measure what it set out to measure? (e.g., controlling variables like age and smoking status when testing a new diet).
Reliability: If the study were repeated by someone else, would they get the same results? Large sample sizes improve reliability.

5. Using Mathematics in Risk

In your exam, you may be asked to interpret statistical tests or calculate values. You should be familiar with:
The Correlation Coefficient: A value (represented by \(r\)) that tells us how strong a correlation is. A value of \(+1\) is a perfect positive correlation, and \(-1\) is a perfect negative correlation.
The Student’s t-test: Used to see if the means (averages) of two groups are significantly different from each other.
The Chi-squared test: Used to see if the observed results are significantly different from the expected results.

Statistical Significance
Scientists usually look for a p-value of less than \(0.05\) (\(p < 0.05\)). This means there is a less than \(5\%\) probability that the results happened by pure chance. If the probability is very low, we say the results are statistically significant.

Key Takeaways

Risk perception is subjective; humans underestimate voluntary/familiar risks and overestimate involuntary/dreaded ones.
Correlation is a link between two variables, but it does not prove that one causes the other.
Causation requires a proven biological mechanism.
• Good studies have large, representative samples and control for other variables to ensure validity and reliability.
• Always look for statistical significance to decide if a result is meaningful or just down to luck.

Common Mistake to Avoid: Don't use the word "proved" too easily in your answers. Instead, say the "data suggests a correlation" or "there is evidence for a causal link."