Welcome to the World of Sampling!
Hello there! As you prepare for your HKICPA QP exams, you’ll find that Business Economics and Statistical Analysis are key tools for any professional accountant. In this chapter, we are going to look at Sampling Techniques. Don’t worry if math isn't your favorite subject! We are going to break this down into simple, logical steps that make sense in a real-world business context.
Why do we need sampling? Imagine you are an auditor and a client has 1,000,000 sales invoices. You can't possibly check every single one—it would take years! Instead, you select a smaller group (a sample) to represent the whole group (the population). If the sample looks good, you can be reasonably confident the whole group is good too. Let’s dive in!
1. Key Concepts: Population vs. Sample
Before we learn the "how," we need to understand the "what."
• Population (N): The entire collection of items or people you want to study. Example: All registered companies in Hong Kong.
• Sample (n): A smaller portion of the population that you actually examine. Example: 500 companies selected from the registry.
• Census: When you study every single member of a population. (Expensive and time-consuming!)
• Parameter: A numerical characteristic of a population (e.g., the average profit of all HK companies).
• Statistic: A numerical characteristic of a sample (e.g., the average profit of the 500 companies you picked).
Memory Trick:
Population = Parameter (Both start with P)
Sample = Statistic (Both start with S)
Quick Review: We use a Sample Statistic to estimate a Population Parameter because it saves time and money.
2. Probability Sampling Techniques
In Probability Sampling, every member of the population has a known, non-zero chance of being selected. This is the "gold standard" for statistics because it avoids bias.
A. Simple Random Sampling (SRS)
This is like a lottery. Every item has an equal chance of being picked. You might use a random number generator to pick invoice numbers.
Example: Putting 100 names in a hat and pulling out 10.
B. Systematic Sampling
You pick a starting point at random and then select every \(k^{th}\) item.
The formula for the interval is: \(k = \frac{N}{n}\)
(Where \(N\) is population size and \(n\) is sample size).
Example: If you have 1,000 invoices and want a sample of 50, you calculate \(1,000 / 50 = 20\). You pick a random start between 1 and 20, then pick every 20th invoice after that.
C. Stratified Random Sampling
You divide the population into groups called strata based on a characteristic (like company size: Small, Medium, Large). Then, you take a random sample from each group.
Why use it? It ensures that important sub-groups are represented. You wouldn't want to accidentally pick only "Small" companies if "Large" companies make up most of the economy!
D. Cluster Sampling
The population is divided into groups called clusters (often based on geography). You randomly select a few entire clusters and study everyone inside them.
Analogy: If you want to study students in Hong Kong, you randomly pick 5 schools (the clusters) and survey every student in those 5 schools.
Key Takeaway: Probability sampling is "fair" and allows us to use mathematical formulas to calculate how accurate our results are.
3. Non-Probability Sampling Techniques
Sometimes, we can't use random selection because it's too difficult or expensive. In Non-Probability Sampling, the chance of being selected is unknown. Be careful: These methods are more likely to be biased!
A. Convenience Sampling
Choosing people or items that are easiest to reach.
Example: Standing outside an MTR station and asking whoever walks by. It's fast, but your sample might only represent people who use that specific station at that specific time.
B. Quota Sampling
Similar to stratified sampling, but not random. You are told to find a specific number of people (a quota). Once you hit the number, you stop.
Example: "Find 20 male shoppers and 20 female shoppers." Once the researcher finds 20 men, they ignore all other men, even if they have useful data.
C. Judgmental (Purposive) Sampling
The researcher uses their own "expert judgment" to choose who to include.
Example: An auditor specifically picking high-value invoices or "suspicious-looking" items to test. This is useful for finding errors but doesn't represent the average invoice.
D. Snowball Sampling
You find one person, and they refer you to others. This "rolls" like a snowball.
Example: Trying to interview CEOs of tech startups. You find one, and they introduce you to three of their friends who are also CEOs.
Quick Review: Non-probability methods are great for exploratory research but are not good for making broad, scientific conclusions about a whole population.
4. Sampling vs. Non-Sampling Errors
No sample is perfect. There will always be some difference between the sample result and the true population result. We call this difference Error.
Sampling Error
This is the natural "gap" between a sample and the population. It happens because a sample is just a piece of the whole. Even if you do everything perfectly, the sample mean will likely be slightly different from the population mean.
Did you know? You can reduce sampling error by increasing the sample size. The more people you ask, the more accurate you'll be!
Non-Sampling Error
These are "human mistakes" or flaws in the process. These are much more dangerous because they can't be fixed by just adding more people to the sample.
• Selection Bias: Your sampling method excluded certain people (e.g., an online survey excludes people without internet).
• Non-Response Bias: People refuse to answer your survey.
• Measurement Error: Poorly worded questions or a broken scale.
• Data Entry Error: Typing "100" when the answer was "10".
Common Mistake to Avoid: Don't assume that a large sample size fixes everything. If your survey question is confusing (non-sampling error), asking 1,000,000 people will just give you 1,000,000 confused answers!
Summary Takeaway:
1. Sampling Error is natural and expected (the "cost" of sampling).
2. Non-Sampling Error is avoidable and results from poor planning or execution.
5. Final Summary Checklist
Before you move on, make sure you can answer these:
• Can I explain the difference between a Parameter and a Statistic?
• Do I know the difference between Stratified (some from all groups) and Cluster (all from some groups) sampling?
• Can I identify a Systematic sampling interval using \(k = N/n\)?
• Do I understand that Non-Sampling Errors can happen even in a Census?
You've got this! Sampling is all about being a detective—choosing the right evidence to get the most accurate picture of the truth. Keep practicing those definitions!