Introduction to Choosing and Evaluating Sampling Methods
Imagine you want to know the favorite music genre of every teenager in the UK. You couldn't possibly ask all of them—it would take years and cost a fortune! Instead, we use sampling. But not all samples are created equal. In this chapter, we will learn how to choose the right "spoonful" from the "pot" of data so that our results are accurate and fair. This is a core part of the Statistical Enquiry Cycle (SEC), where we plan how to collect data to avoid bias.
The "Big Five" Sampling Methods
The Edexcel syllabus requires you to understand and evaluate five specific ways of picking your sample. Each has its own "personality," strengths, and weaknesses.
1. Simple Random Sampling (SRS)
This is the "gold standard" of sampling. In a simple random sample of size \(n\), every member of the population has an equal chance of being included, and every possible subset of size \(n\) is equally likely to be the sample.
How to do it: Assign every member of the population a number and use a random number table or a calculator to pick your winners.
Unrestricted vs. Simple: Unrestricted means sampling with replacement (you could pick the same person twice), while Simple usually refers to sampling without replacement.
Pros: It is completely unbiased in theory.
Cons: You need a full list of the population (a sampling frame), which can be hard to get. It can also be expensive if the people picked are spread all over the country.
2. Systematic Sampling
Think of this as the "mathematical shortcut." You pick a starting point at random and then take every \(k^{th}\) member (e.g., every 10th person on a list).
Pros: Very simple to carry out and spreads the sample evenly across the list.
Cons: If the list has a hidden pattern (e.g., every 10th house is a corner house with a larger garden), your data will be biased.
3. Cluster Sampling
Sometimes the population is already divided into groups, like classrooms in a school or streets in a town. In Cluster Sampling, you randomly pick a few of these groups (clusters) and sample everyone inside them.
Pros: Much cheaper and quicker for large geographical areas.
Cons: People in the same cluster might be very similar to each other, which might not represent the whole population well.
4. Judgmental Sampling
This is a non-random method. The researcher uses their own "judgment" to choose people they think are representative of the population.
Pros: Good for very small-scale research or when you need specific experts.
Cons: Highly prone to researcher bias. It is not scientifically reliable for making big claims about a population.
5. Snowball Sampling
Used when the population is "hidden" or hard to reach (e.g., people with a rare hobby). You find one person, interview them, and then ask them to "nominate" friends to be interviewed next.
Pros: Allows you to reach people who wouldn't normally answer a survey.
Cons: The sample isn't random at all; people tend to nominate friends who think just like them!
Quick Review: Which method is best if you don't have a list of the population but need to reach people in a specific subculture? (Answer: Snowball Sampling!)
Stratification: The "Layer Cake" Approach
Sometimes, we want to make sure specific groups (like different age brackets or genders) are represented correctly. We call these groups strata. We stratify before we start sampling.
Proportional Stratification
The number of people in your sample for each group matches the percentage in the real population. If \(60\%\) of a school is female, then \(60\%\) of your sample should be female.
Formula: \(\text{Number in sample} = \frac{\text{Number in stratum}}{\text{Number in population}} \times \text{Total sample size}\)
Disproportional Stratification
Sometimes, a group is so small that a proportional sample would only include one or two people—not enough to learn anything! In disproportional sampling, you might over-sample that small group to ensure you have enough data to analyze them properly.
Key Takeaway: Stratification ensures no group is left out, making the sample more representative than a basic random sample.
Choosing the Right Method for the Context
In your exam, you will be asked to choose or evaluate a method based on a specific scenario. Here is a guide:
Market Research
Often uses Judgmental or stratified methods. You need to reach specific types of consumers quickly to see if a product will sell.
Exit Polls
When people leave a voting station, we often use Systematic sampling (e.g., asking every 10th voter). This avoids the interviewer just picking people who "look friendly."
Quality Assurance (QA)
In a factory, you might use Systematic sampling on the assembly line to check for defects. It’s practical and ensures the whole day's production is checked.
Evaluating Constraints and Bias
When evaluating a sampling plan, always check for practical constraints and bias. Even a perfect plan can fail if:
- Time/Cost: Is the method too expensive? (SRS is often costly).
- Access: Can you actually reach the people? (Snowball is best for hard-to-reach groups).
- Leading Questions: Even with a perfect random sample, if your questions are "loaded" (e.g., "Don't you agree that..."), your data is biased.
- Sample Size: If \(n\) is too small, the results won't be reliable.
Don't forget: A conclusion from a sample is never definite. Because we didn't ask everyone, there is always a chance the sample is slightly different from the true population.
Common Mistakes to Avoid
- Confusing Cluster and Stratified: In Stratified, you take a few people from all groups. In Cluster, you take all people from a few groups.
- Ignoring the Sampling Frame: You can't do a Simple Random Sample if you don't have a list of every member of the population!
- Replacement Errors: Remember that "without replacement" (standard SRS) is different from "with replacement" (unrestricted). In most real-world cases, we don't want to interview the same person twice!
Summary Checklist
Before you answer an exam question, ask yourself:
1. Do I have a list of the population? (If yes, think SRS or Systematic).
2. Are there specific groups I must represent? (Think Stratification).
3. Is the population hard to find? (Think Snowball).
4. Is it a geographical problem? (Think Cluster).
5. Is there a risk of bias? (Always check for this!).