Welcome to Data Collection: The Foundation of Success

Ever heard the phrase "Garbage in, garbage out"? In AP Statistics, that is our golden rule! You can have the most advanced calculator in the world, but if your data collection is messy or biased, your results won't mean a thing. In this chapter, we are going to learn how to choose the right way to gather data and—more importantly for the AP Exam—how to justify those choices using statistical reasoning. Whether you're preparing for the digital Bluebook exam or just trying to survive your next quiz, these notes have you covered!

Note: This chapter focuses on Practice 2: Collect Data from the updated AP Statistics curriculum (effective Fall 2026).

1. The Big Picture: Observation vs. Experiment

Before we pick a method, we have to decide what kind of "investigative question" we are answering. (For more on questions, see our chapter on "Formulating an investigative question").

Observational Studies

In an observational study, researchers observe individuals and measure variables of interest but do not attempt to influence the responses. You are a fly on the wall.
Best for: Identifying correlations or describing a population.

Experiments

In an experiment, researchers deliberately impose some treatment on individuals to measure their responses.
Best for: Determining cause-and-effect (causation).
The Secret Sauce: For a study to be a valid experiment, it must include random assignment of treatments. Without random assignment, you cannot claim that one thing caused another!

Quick Review: If the goal is to show that "Medicine A causes a faster recovery," you need an experiment. If you just want to know "What percentage of people take Medicine A," an observational study (like a survey) is perfect.

2. How to Pick Your Sample (Sampling Methods)

When we want to know about a population (the whole group), we usually only look at a sample (a part of the group). Here are the four "valid" random sampling methods you need to know for Practice 2.A.

1. Simple Random Sample (SRS): Every group of size \( n \) has an equal chance of being chosen.
Analogy: Putting everyone's name in a giant hat, shaking it well, and pulling out 10 names.

2. Stratified Random Sample: Divide the population into groups of similar individuals (called strata) and then take a separate SRS from each group.
Example: If you want to know how students feel about school lunch, you might split the school into "9th grade," "10th grade," etc., and sample 20 students from every grade. This ensures every grade is represented.

3. Cluster Sample: Divide the population into groups that are located near each other (called clusters). Randomly pick a few clusters and survey everyone in those selected groups.
Example: To survey a city, randomly pick 5 apartment buildings and interview every single person living in those 5 buildings.

4. Systematic Random Sample: Choose a random starting point and then pick every \( k^{th} \) person.
Example: Stand at the school door and survey every 10th person who walks in (after starting at a randomly chosen person between 1 and 10).

Key Takeaway: If a question asks you to justify a sampling method, explain how it reduces variability (like Stratified) or makes the process more practical (like Cluster or Systematic).

3. Identifying Potential Problems (What Could Go Wrong?)

Even with good intentions, data collection can go off the rails. The AP Exam loves to ask you to identify these errors (Practice 2.D).

  • Undercoverage: Some members of the population are left out of the process of choosing the sample. (Example: A phone survey that ignores people without landlines.)
  • Nonresponse: An individual chosen for the sample can't be contacted or refuses to participate. (Example: You mail 100 surveys but only 5 people mail them back.)
  • Response Bias: A systematic pattern of incorrect responses. This could be caused by the wording of the question, the interviewer's behavior, or people lying about embarrassing topics.

Common Mistake to Avoid: Don't confuse "Voluntary Response" with "Nonresponse." Voluntary Response is when anyone can choose to join (like an open internet poll). Nonresponse is when you've already picked someone to be in your sample, but they don't answer.

4. Designing a Valid Experiment

If you are asked to "justify an appropriate method" for an experiment (Practice 2.B), you must address these four pillars:

  1. Comparison: Use a design that compares two or more treatments (one might be a control group).
  2. Random Assignment: Use a chance process to assign experimental units to treatments. This helps balance out variables we aren't measuring!
  3. Control: Keep other variables the same for all groups so they don't mess up the results.
  4. Replication: Use enough experimental units in each group so that any differences in the effects of the treatments can be distinguished from chance variations.
Advanced Experimental Tools:

Blocking: If you think a certain characteristic (like age or gender) will affect the response to the treatment, group your subjects by that characteristic first (forming "blocks"), then randomly assign treatments within each block.
Note: Blocking is to experiments what stratifying is to sampling!

Blinding:
- Single-blind: The subject doesn't know which treatment they are getting.
- Double-blind: Neither the subject nor the researcher interacting with them knows who is getting what. This prevents the researcher from accidentally influencing the results.

5. Writing Your Justification (Exam Tips)

On Section II (Free-Response) of the digital AP Exam, you will often be asked to justify why a certain method is appropriate. Use this checklist:

  • Context: Always use the names of the variables and the people/items in the study. Don't just say "the subjects," say "the tomato plants" or "the student athletes."
  • Randomness: Explicitly state that random selection (for sampling) or random assignment (for experiments) was used to avoid bias or balance variables.
  • Causation: If the question asks if we can conclude that \( X \) caused \( Y \), your justification must mention that the treatments were randomly assigned.
  • Scope of Inference:
    - If we randomly sampled from the population, we can generalize our results to that population.
    - If we randomly assigned treatments, we can make a causal claim.

Don't worry if this seems like a lot to remember! Just keep asking yourself: "Was it random?" and "Who is being studied?" If you can answer those two questions, you're 90% of the way to a perfect score on Practice 2!

Looking for the next step? Check out the chapter on "Choosing an inference method and verifying conditions" to see what to do once your data is collected!