Welcome to the World of Data Collection!

Before a statistician can create fancy graphs or calculate averages, they need something very important: data. But you can’t just grab any old numbers! How you collect your data determines whether your results are useful or complete nonsense. In this chapter, we’ll explore the different ways to gather information, how to design a great questionnaire, and how to set up experiments that actually work.

1. Where Does Data Come From? (Sources)

There are several ways to gather data, depending on what you want to find out. Here are the main methods mentioned in your syllabus:

  • Experimental: You actively change something to see what happens.
    • Laboratory: Done in a controlled environment (like a science lab).
    • Field: Done in a real-life setting (like a classroom or a park).
    • Natural: You observe a change that happens on its own (like comparing data before and after a new law is passed).
  • Simulation: Sometimes, a real experiment is too dangerous, expensive, or slow. We use a simulation (often on a computer) to mimic real life. Example: Simulating traffic flow to design a new junction.
  • Questionnaires: Asking people questions directly. This can be done on paper, online, or via an interview.
  • Observation: Watching and recording what happens without interfering. Example: Tallying the types of cars that pass a school.
  • Reference and Census: Using data that already exists. A census is data collected from every single member of a population.

Quick Tip: If you use data someone else collected, it is called secondary data. You must always acknowledge your sources (say where you got it from) to stay ethical and professional!

2. Designing Questionnaires

A questionnaire might seem easy, but asking the wrong questions can lead to bias (results that aren't fair or accurate). Here is what you need to know:

Open vs. Closed Questions

  • Open Questions: These allow the person to answer however they like.
    Example: "What do you think of school dinners?"
    Pros: Lots of detail. Cons: Very hard to turn into a graph.
  • Closed Questions: These give fixed options to choose from.
    Example: "Are you happy with school dinners? (Yes/No)" or a scale of \(1\) to \(5\).
    Pros: Very easy to analyse and put into a chart. Cons: People might feel their true opinion isn't an option.

The "Golden Rules" of Questions

  1. Avoid Leading Questions: Don't push people toward a specific answer. Instead of saying "Don't you agree that homework is boring?", ask "How do you feel about your homework?"
  2. Avoid Ambiguity: Make sure there is no confusion. Don't ask "Do you exercise often?" because "often" means different things to different people. Use "How many times a week do you exercise?" instead.
  3. Mutually Exclusive Boxes: Ensure options don't overlap.
    Bad example: \(0-10\) hours, \(10-20\) hours. (Where does someone who does exactly \(10\) go?)
    Good example: \(1-10\) hours, \(11-20\) hours.

Did you know? Sensitive subjects (like illegal activity or personal health) can make people lie. This causes bias. To help with this, researchers sometimes use the Random Response technique (Higher Tier only), which uses a random event like a coin flip to keep the person's individual answer secret while still allowing the researcher to calculate an overall percentage.

3. Experiments and Design (Higher Tier Focus)

If you are taking the Higher Tier exam, you need to know a bit more about how to make an experiment "fair."

  • Control Groups: This is a group that does not receive the treatment. We compare them to the group that did receive the treatment to see if there is a real difference.
  • Matched Pairs: To make a test fair, you pair up participants who are very similar (same age, same fitness level, etc.). One person in the pair gets the treatment, and the other doesn't.
  • Extraneous Variables: These are "nuisance" variables that might mess up your results. For example, if you are testing if a plant grows better with a new fertiliser, you must make sure all plants get the same amount of sunlight. Sunlight is an extraneous variable you must control.

4. Reliability and Validity

These two words sound similar, but they mean different things in Statistics:

Reliability: If you did the experiment again, would you get the same result? If the answer is yes, your data is reliable. Using a larger sample size usually makes results more reliable.

Validity: Does the data actually measure what it's supposed to measure? If you want to know how fit someone is, measuring their height is not a valid method!

5. Cleaning Data

Before you start your calculations, you need to "clean" your data. This involves:

  • Identifying missing responses (people who skipped a question).
  • Spotting errors (like someone writing "150" for their age when they meant "15").
  • Dealing with outliers (values that are very different from the rest).

Common Mistake: Many students forget that "cleaning" isn't just about deleting weird data—it's about deciding whether that data is a genuine extreme value or just a mistake.

Key Takeaways for Revision

1. Choose the right source: Use experiments for cause-and-effect, and observations or questionnaires for opinions/behaviours.
2. Watch your language: In questionnaires, avoid leading questions and ensure tick-boxes don't overlap.
3. Stay Valid: Ensure you are actually measuring what you claim to be measuring.
4. Identify Bias: Always look for reasons why the data might be "tilted" or unfair, such as a biased source or sensitive questions.

Note: For more on how to pick who to ask, check out the chapter on "Population and Sampling". For more on the different categories of numbers you collect, see "Types of Data".