Planning and Data Collection: Primary and Secondary Data
Welcome to your study guide on Primary and Secondary Data Collection Methods! Whenever you carry out a statistical investigation in your CCEA GCSE Statistics course, the very first thing you need to decide is: "Where will I get my information from?"
Don't worry if statistics sometimes feels overwhelming. In this chapter, we will break down the two main types of data, look at how each is collected, and explore their pros and cons. By the end of these notes, you will know exactly how to pick the right method for any exam question!
1. What is the Difference? (Core Definitions)
Every piece of data used in an investigation falls into one of two categories based on who collected it and why it was collected.
Primary Data:
Data that is collected directly by or on behalf of the person carrying out the research from an original source for a specific, stated purpose.
Analogy: Making a fresh pizza from scratch. You pick every topping yourself, and it is made specifically for your dinner.
Secondary Data:
Data that has already been collected by someone else or by an external organisation and is now being re-used for a new or existing investigation.
Analogy: Buying a frozen pizza from the supermarket. It is ready immediately and saves you time, but you cannot change the toppings that are already on it.
Memory Trick:
• Primary = Personal (you collected it yourself for your Purpose).
• Secondary = Second-hand (someone else collected it first).
Key Takeaway: If you (the investigator) gathered the data directly for your own project, it is primary. If you found or downloaded data that someone else already gathered, it is secondary.
2. Methods of Collecting Primary Data
When you decide to collect primary data, you have several methods to choose from depending on what you are investigating:
A. Surveys and Questionnaires
Surveys are used to collect opinions, facts, or numbers from people. They can be carried out in four main ways:
• Face-to-face interviews: Asking questions in person. This gives high response rates and lets you explain questions, but it takes a lot of time and can be expensive.
• Postal surveys: Sending written questionnaires in the post. This can reach a wide area, but response rates are often very low.
• Telephone interviews: Calling people directly to ask questions. Faster than face-to-face, but people may hang up or not answer unknown numbers.
• Online / Web-based questionnaires: Creating a survey on a website. Very quick and cheap to distribute widely, but only people with internet access can take part.
Design Tip: When writing questionnaire questions, always make sure questions are unbiased, clear, and avoid leading the respondent toward a particular answer.
B. Direct Observation
Watching and recording events or behaviour directly as they happen without interfering.
• Examples: Using a tally chart to count the number of cars passing a school gate, or using automated data-logging counters to monitor foot traffic in a shopping centre.
C. Controlled Experiments and Clinical Trials
Setting up a controlled test to establish cause-and-effect relationships.
• Involves an explanatory variable (the factor you change) and a response variable (the outcome you measure).
• Often uses a treatment group (receives the test item/medicine) and a control group (does not receive it, acting as a baseline for comparison).
D. Direct Measurement and Data Logging
Using measuring instruments or electronic sensors to collect objective numerical measurements.
• Examples: Using a stopwatch to record running times, a ruler to measure plant height, or automated electronic sensors to record continuous temperature readings over time.
Key Takeaway: Primary data collection includes surveys (face-to-face, postal, phone, online), direct observation, controlled experiments, and direct physical measurement.
3. Sources of Secondary Data
If you do not have the time or resources to collect your own data, you can look for existing datasets from reliable sources:
• Official Government Statistics: Highly reliable, large-scale data published by national bodies such as the Northern Ireland Statistics and Research Agency (NISRA), the Office for National Statistics (ONS), and the Department for Education.
• Published Databases and Repositories: National Census records, historical weather/meteorological archives, public online databases, and academic research papers.
• Commercial and Media Sources: Market research reports, published books, newspaper articles, company annual financial reports, and business directories.
Did You Know? The national Census collects information about the entire population every 10 years. It provides one of the largest and most comprehensive secondary datasets available for statistical research in the UK!
Key Takeaway: Reliable secondary data can come from government bodies like NISRA and ONS, national census archives, meteorological records, and published commercial reports.
4. Comparing Advantages and Disadvantages
In your CCEA exam, you will frequently be asked to evaluate whether a researcher should use primary or secondary data. Use the detailed points below to give strong, specific reasons.
Primary Data: Evaluation
Advantages:
• Exact fit: Collected for the exact, specific purpose of the investigation.
• Full control: You know exactly how the data was collected, how the sample was chosen, and how reliable it is.
• Up to date: The data is collected in real time, so it is current.
• Quality control: You control operational definitions, units of measurement, and accuracy.
Disadvantages:
• Time-consuming: Takes a lot of time to design, pilot test, and carry out.
• Expensive: Can require significant money for staff, travel, postage, or equipment.
• Limited scale: Sample sizes collected by one person or a small team are often small or restricted to one local area.
• Non-response bias: People may refuse to take part or fail to return postal surveys, leaving you with an unrepresentative sample.
Secondary Data: Evaluation
Advantages:
• Quick and convenient: Readily accessible immediately without waiting for fieldwork.
• Low cost: Usually free or very cheap to obtain.
• Massive datasets: Gives access to huge, nationwide samples (like Census data) that would be impossible for an individual student or researcher to gather.
• Historical comparisons: Allows you to study trends and patterns over many years or decades.
Disadvantages:
• Poor fit: It was collected for a different purpose, so the questions or categories may not match your hypothesis.
• Unknown collection methods: You may not know if there was bias, sampling errors, or poor methodology in the original collection.
• Out of date: The information may be old or obsolete.
• Inconvenient groupings: Data may already be grouped into class intervals that do not suit your calculations.
5. Practical Considerations & Data Cleaning
Before and after collecting data, statisticians must manage real-world limitations and ensure data quality.
Practical Constraints on Data Collection
When planning an investigation, researchers must balance five key constraints:
1. Time: Deadlines for finishing the research.
2. Cost: The financial budget available for staff and equipment.
3. Ethical Factors & Data Protection: Ensuring participant privacy, anonymity, and informed consent.
4. Health & Safety: Making sure data collectors are not placed in dangerous situations.
5. Ease of Access: Having access to the target population or a suitable sampling frame.
Data Cleaning
Once data has been collected, it is rarely perfect. Data cleaning is the process of checking raw data before analysis begins to fix or remove problems:
• Identifying Anomalies and Outliers: Spotting extreme values that do not fit the rest of the pattern (for example, a human height recorded as \(240\text{ cm}\) or an age written as \(-5\)).
• Handling Missing Data: Deciding how to deal with blank spaces where participants skipped questions.
• Correcting Transcription Errors: Fixing typing mistakes made when entering paper data into a computer spreadsheet.
Key Takeaway: Always consider practical constraints (cost, time, ethics) during planning, and clean your data to remove errors and anomalies before calculating statistics.
6. Top Exam Pitfalls & Examiner Tips
Make sure you avoid these common mistakes highlighted by examiners:
Pitfall 1: Confusing the internet with the type of data
• Mistake: Thinking that anything downloaded from the internet is automatically secondary data, or that online surveys are secondary.
• Correct Rule: If you write and host an online questionnaire to gather responses, it is primary data. If you download a published spreadsheet from NISRA or the ONS website, it is secondary data.
Pitfall 2: Giving vague, one-word answers
• Mistake: Writing "It is easy" or "It is hard" as an advantage or disadvantage.
• Correct Rule: Always be specific! Write: "Secondary data is cheaper and faster to access than conducting an original survey" or "Secondary data might not match the exact variables needed for the hypothesis."
Pitfall 3: Assuming secondary data is always 100% accurate
• Mistake: Believing secondary sources have no mistakes.
• Correct Rule: Always evaluate secondary data critically. The original author might have used biased questions, or the data may now be out of date.
Pitfall 4: Forgetting non-response bias
• Mistake: Ignoring what happens when people refuse to answer surveys.
• Correct Rule: Remember that people who choose to respond to voluntary postal or online surveys often have stronger opinions than those who ignore them, which can bias the results.
Quick Review Quiz Checklist
Can you answer these key questions from memory?
1. If a researcher counts traffic using a tally chart at a crossroads, is this primary or secondary data?
2. What are two advantages of using Census data from NISRA/ONS rather than carrying out your own survey?
3. Name two reasons why a researcher must carry out data cleaning before analysing raw data.
4. Why is a postal questionnaire vulnerable to non-response bias?