Welcome to Handling Data and Data Collection!
Have you ever wondered how weather forecasters predict rain, how video game developers balance difficulty, or how streaming services know which songs to recommend? The answer is Data! In this chapter, we will explore how mathematicians and scientists collect, organize, and make sense of information from the world around us. Don't worry if maths hasn't always felt easy — this topic is very logical, practical, and connected to everyday life.
1. The Handling Data Cycle
Whenever we want to investigate a real-world question using mathematics, we follow a four-stage process known as the Handling Data Cycle. Think of it like a continuous loop: once you answer one question, it often leads to new ones!
Stage 1: Specify the Problem and Plan
First, decide what you want to find out. We usually start with a clear question or a hypothesis (an idea or prediction that can be tested).
Example: "Students who eat breakfast score higher in morning tests than those who do not."
You also decide what data you need to collect and how you will get it.
Stage 2: Collect the Data
Next, gather the facts, numbers, or responses using surveys, experiments, observations, or existing records. Ensuring the data is accurate and fair is crucial at this step.
Stage 3: Process and Represent the Data
Turn your raw numbers into clear information. This involves organizing data into tables, calculating averages (mean, median, mode) and spread (range), and drawing charts or diagrams (such as bar charts, pie charts, or histograms).
Stage 4: Interpret and Discuss the Results
Look at your graphs and calculations to answer your original question. Did your findings support your hypothesis? What limitations did your investigation have, and what could you improve next time?
Memory Aid — Remember the 4 Steps:
Please Collect Pie Ingredients → Plan, Collect, Process, Interpret.
Key Takeaway: The data cycle always starts with a plan/hypothesis and finishes with an interpretation of the results to draw a conclusion.
2. Classifying Types of Data
Before you collect any data, you need to know what kind of data you are dealing with. Data can be split into different categories.
A. Qualitative vs. Quantitative Data
Qualitative Data:
This is non-numerical data described in words or labels. It describes qualities or characteristics.
Examples: Eye colour (blue, brown, green), favourite subject, car brand, or flavours of crisps.
Quantitative Data:
This is numerical data (data that consists of numbers/quantities) that can be counted or measured.
Examples: Number of pets (\(3\)), shoe size (\(6\)), height (\(165\text{ cm}\)), or time taken to run \(100\text{ m}\) (\(14.2\text{ s}\)).
B. Discrete vs. Continuous Data (Splitting Quantitative Data)
Quantitative data is divided into two very important types:
1. Discrete Data:
Data that can only take specific, distinct values. It is usually counted in whole units, with nothing in between.
Key test: Ask yourself, "Can I count this in separate units?" If yes, it is discrete.
Examples: Number of goals scored in a match (\(0, 1, 2, 3\)), number of siblings, or shoe size (you can have size \(5\) or \(5.5\), but not size \(5.234\)).
2. Continuous Data:
Data that can take any numerical value within a given range. It is measured on a scale and can be broken down into smaller and smaller fractions or decimals depending on the accuracy of your measuring tool.
Key test: Ask yourself, "Is this measured using a tool (like a ruler, scale, or stopwatch)?" If yes, it is continuous.
Examples: Height (\(1.72\text{ m}\)), weight (\(54.65\text{ kg}\)), temperature (\(21.4^\circ\text{C}\)), or time taken to complete a test (\(42.3\text{ minutes}\)).
Did You Know? Time is always continuous because between any two points in time (e.g., \(10\text{ seconds}\) and \(11\text{ seconds}\)), there are infinite smaller fractions of a second!
C. Primary vs. Secondary Data
Primary Data:
Data that you (or your team) collect yourself for your specific investigation.
Methods: Handing out your own questionnaire, conducting an experiment, or carrying out a traffic count.
Advantages: Up-to-date, relevant to your exact question, you know how reliable it is.
Disadvantages: Can take a lot of time, effort, and money to collect.
Secondary Data:
Data that has already been collected by someone else for another purpose.
Sources: The internet, government census databases, newspapers, weather station archives.
Advantages: Quick, easy, and cheap to find; allows access to large datasets.
Disadvantages: May be out of date, may not answer your exact question, or could be inaccurate/biased.
Key Takeaway:
• Qualitative = Words.
• Quantitative = Numbers.
• Discrete = Counted (distinct steps).
• Continuous = Measured (any value on a scale).
• Primary = Collected by you; Secondary = Collected by someone else.
3. Data Collection Tools: Questionnaires and Tally Sheets
To collect reliable primary data, you must design effective tools.
A. Designing Good Questionnaires
In your exam, you will often be asked to critique a poorly designed question or write a better version. Here are the core rules of a good survey question:
1. Provide clear, non-overlapping response boxes (options):
Poor question: How many books did you read last month?
[ ] \(0 - 2\) [ ] \(2 - 4\) [ ] \(4 - 6\)
Problem: If someone read \(2\) books, which box do they tick? The boxes overlap!
Better question:
[ ] \(0\) [ ] \(1 - 2\) [ ] \(3 - 4\) [ ] \(5\text{ or more}\)
2. Cover all possible options (Exhaustive):
Ensure no one is left without a valid option. Including categories like "\(0\)" or "More than..." prevents people from being excluded.
3. Avoid vague language — include a specific timeframe:
Poor question: "Do you exercise often? Yes [ ] No [ ]"
Problem: What does "often" mean? Every day? Once a week? Once a month?
Better question: "How many times did you exercise for at least \(20\text{ minutes}\) in the past \(7\text{ days}\)?"
[ ] \(0\) times [ ] \(1 - 2\) times [ ] \(3 - 4\) times [ ] \(5\text{ or more}\) times
4. Avoid Leading or Biased Questions:
A question should not push the respondent toward a certain answer.
Poor question: "Don't you agree that homework is boring?" (Leading)
Better question: "To what extent do you enjoy completing homework?"
B. Tally Charts and Frequency Tables
A tally chart is a quick way to record data as it happens in real time. We group tally marks in bundles of five (four vertical strokes and one diagonal stroke across: \(||||\hspace{-1.1ex}/\,\)) to make counting easy.
A frequency table converts these tallies into clear numbers representing the total count (frequency) for each category.
Example: A survey of favourite fruit colours among \(20\) students:
• Red: Tally = \(||||\hspace{-1.1ex}/\,\) \(|||\) → Frequency = \(8\)
• Yellow: Tally = \(||||\hspace{-1.1ex}/\,\) → Frequency = \(5\)
• Green: Tally = \(||||\hspace{-1.1ex}/\,\) \(||\) → Frequency = \(7\)
Total Frequency = \(8 + 5 + 7 = 20\).
Key Takeaway: Good survey questions are fair, specific, have fixed timeframes, and offer tick boxes that do not overlap.
4. Populations, Samples, and Bias
When you carry out an investigation, you need to decide who or what to study.
A. Population vs. Sample
Population:
The entire group of individuals, animals, or items that you want to find out about.
Examples: All students in a school, all voters in Northern Ireland, or every lightbulb made in a factory.
Census:
A survey that collects information from every single member of the entire population.
Advantage: Completely accurate representation of the population.
Disadvantage: Very expensive, time-consuming, and often impossible for large populations.
Sample:
A smaller selection of items or people chosen from the population to represent the whole group.
Advantage: Faster, cheaper, and more practical than testing everyone.
Disadvantage: Might not be completely representative if not chosen carefully.
B. Understanding and Avoiding Bias
A sample is biased if it systematically favours certain outcomes or does not fairly represent the entire population.
Example of a Biased Sample:
Asking \(50\) people outside a gym about their favourite weekend hobby. This is biased because gym visitors are far more likely to prefer fitness and sports than the general public.
How to avoid bias:
1. Use Random Sampling: Give every member of the population an equal chance of being selected (e.g., assigning each student a number and picking numbers using a random number generator).
2. Use an Adequate Sample Size: A sample of \(5\) people is too small to represent a school of \(1000\). A larger sample size generally gives more reliable results.
3. Sample at different times and locations: Ensure varied demographics are represented.
Key Takeaway: A sample saves time and money, but it must be large enough and randomly chosen to avoid bias.
5. Common Mistakes to Avoid in Exams
• Confusing Discrete and Continuous: Remember, if it's measured with an instrument (like a stopwatch or ruler), it is continuous, even if rounded to the nearest whole number!
• Overlapping Intervals: Writing options like "\(10 - 20\)" and "\(20 - 30\)". Always write "\(10 - 19\)" and "\(20 - 29\)" or use inequality notation such as \(10 \le x < 20\) and \(20 \le x < 30\).
• Forgetting "Zero" or "Other": Always include options for non-participants (e.g., "\(0\) times") or alternative options.
• Saying "Ask more people" without explaining why: In exam evaluation questions, explain that a larger sample size reduces the effect of anomalies and makes the sample more representative.