Welcome to Types of Data

Welcome to your study notes for Types of Data! In Statistics, everything begins with data. Before you can draw fancy charts, calculate averages, or make predictions, you need to understand what kind of information you are working with.

Think of data like ingredients in a kitchen: you wouldn't use salt the same way you use sugar! Knowing your data type helps you choose the correct statistical tools, diagrams, and formulas. Don't worry if this seems tricky at first—we will break down every concept step-by-step with simple examples and memory tricks.


1. Qualitative vs Quantitative Data

The very first question to ask about any piece of data is: Is it made of words or numbers?

Qualitative Data (Categorical Data)

Qualitative data is descriptive and non-numerical. It describes qualities, characteristics, or categories.

Examples: Favourite colours (blue, green, red), eye colour, types of pets, car brands, or town names.
Memory Trick: Think of the letter L in Qualitative for Letters and Labels.

Quantitative Data (Numerical Data)

Quantitative data is numerical information that can be counted or measured. It answers questions like "how many?" or "how much?"

Examples: Number of siblings (\(3\)), shoe size (\(6\)), height (\(168\text{ cm}\)), or temperature (\(18.5^\circ\text{C}\)).
Memory Trick: Think of the letter N in Quantitative for Numbers.

Common Mistake to Avoid

Be careful with numbers that are used only as labels! For example, a football player's squad number (e.g., Number 9) or a postcode is actually qualitative data, because you cannot do meaningful maths with it (calculating the average postcode makes no sense!).

Key Takeaway:
Qualitative: Words / Categories (e.g. favourite subject)
Quantitative: Numbers / Quantities (e.g. test score)


2. Discrete vs Continuous Data

If your data is quantitative (numbers), you must classify it further into either discrete or continuous.

Discrete Data

Discrete data can only take specific, distinct values. It cannot take any value in between. Most of the time, discrete data is counted in whole numbers.

Examples:
- Number of students in a classroom (you can have \(24\) or \(25\) students, but never \(24.5\) students).
- Goals scored in a football match (\(0, 1, 2, 3\)).
- Shoe sizes (\(5, 5.5, 6, 6.5\) — even though half-sizes exist, you cannot buy a shoe size of \(5.273\)).

Continuous Data

Continuous data can take any value within a given range. It is not restricted to fixed steps and depends on how accurately you measure it. Continuous data is always measured using an instrument (like a ruler, scale, or stopwatch).

Examples:
- Height of a tree (e.g., \(4\text{ m}\), \(4.2\text{ m}\), \(4.238\text{ m}\)).
- Time taken to run \(100\text{ m}\) (e.g., \(12.4\text{ seconds}\)).
- Mass of an apple (e.g., \(150.25\text{ g}\)).

Analogy: Staircase vs Ramp

Discrete data is like climbing a staircase: you can stand on step \(1\) or step \(2\), but you cannot hover in mid-air between the steps.
Continuous data is like walking up a ramp: you can smoothly stop at any height imaginable along the slope.

Quick Checklist: "Counted or Measured?"

Ask yourself:
1. Did I count it? \(\implies\) Discrete
2. Did I measure it? \(\implies\) Continuous

Key Takeaway:
Discrete: Fixed, separate values (counted).
Continuous: Any value on a continuous scale (measured).


3. Primary vs Secondary Data

When planning an investigation, you must decide where your data will come from.

Primary Data

Primary data is information collected firsthand by you (or your team) specifically for your own investigation.

Methods of collection: Questionnaires, surveys, personal observations, experiments.
Advantages:
- You know exactly how it was collected and how reliable it is.
- It is tailored directly to your specific research question.
- It is up to date.
Disadvantages:
- Can be time-consuming, expensive, and difficult to collect in large amounts.

Secondary Data

Secondary data is information that has already been collected by someone else for a different purpose, which you then use.

Sources: Government census data, websites, weather records, books, newspaper articles.
Advantages:
- Quick, easy, and often free to obtain.
- Gives access to huge sample sizes (like whole-country population data).
Disadvantages:
- Might be out of date or contain errors.
- It might not answer your exact question.
- You do not know if there was bias in how it was gathered.

Key Takeaway:
Primary: Collected by you directly.
Secondary: Collected by others and borrowed by you.


4. Grouped vs Ungrouped Data

Once data is collected, we need to organise it to spot patterns.

Ungrouped Data (Raw Data)

Ungrouped data is simply a list of individual values exactly as they were recorded.

Example: Test marks of \(5\) students: \(12, 18, 14, 12, 19\).
Best for: Small sets of data where you want to see every single value.

Grouped Data

When dealing with large amounts of data, individual lists become overwhelming. Grouped data sorts numbers into class intervals (groups).

Example: Heights of students organized into intervals:
- \(140 \le h < 150\) (Frequency: \(4\))
- \(150 \le h < 160\) (Frequency: \(11\))
- \(160 \le h < 170\) (Frequency: \(7\))
Advantage: Summarises large data sets clearly and makes patterns easy to see.
Disadvantage: You lose the original individual values (you know \(4\) students are between \(140\text{ cm}\) and \(150\text{ cm}\), but you don't know their exact heights anymore!).

Key Takeaway:
Ungrouped: Every individual value is visible.
Grouped: Values are bundled into class intervals for easier summary.


Master Summary Table

Use this simple table as a quick-revision checklist for your exam:

Qualitative: Described by words/labels (e.g. Car colour)
Quantitative: Described by numbers (e.g. Speed of car)
Discrete: Counted numbers with distinct jumps (e.g. Number of passengers)
Continuous: Measured numbers that can take any value (e.g. Journey time)
Primary: Collected directly by the researcher (e.g. Your own survey)
Secondary: Taken from existing sources (e.g. Internet statistics)
Grouped: Organised into class intervals (e.g. \(20 \le x < 30\))
Ungrouped: Individual list of raw numbers (e.g. \(21, 24, 28\))

Exam Tip

In exam questions, always classify data in pairs! For example, if asked to describe the variable "Time spent revising", state both Quantitative and Continuous to gain full marks.