Executive Summary & Difficulty Verdict
The 2023 AP Statistics Free-Response section presented a balanced yet technically rigorous assessment across standard core competencies and non-routine problem solving. With a global mean score of 2.89 and an average score of 0.88/4 on Question 4 (Paired Means Inference), the FRQ section exhibited a moderate-to-challenging difficulty profile. Students demonstrated strong procedural fluency in standard computations such as discrete expected value and least-squares regression arithmetic, but encountered notable friction on experimental precision, paired inference conceptualization, and simulation-based sampling distributions in the Investigative Task.
Where the Marks Were Won and Lost
- Unit 1 (One-Variable Data): High accessibility on boxplot construction (Q1b), but marks were dropped in Q1a and Q1c when students failed to systematically address all four distributional characteristics (C-U-S-S: Center, Unusual features, Shape, Spread) with comparative language and units.
- Unit 3 (Collecting Data): Q2 saw frequent drops in communication points when students omitted essential randomization mechanics (e.g., failing to specify sampling without replacement or neglecting to define the integer range for random number generators).
- Unit 4 (Probability): Discrete calculations in Q3a and Q3c were well handled, but conditional probability in Q3b suffered from the common misstep of assuming independence and multiplying joint events.
- Unit 7 (Inference for Means): The steepest mark drop in the exam occurred on Q4, where students persistently misidentified a paired \(t\)-test as a two-sample \(t\)-test and struggled to assess the normality condition using differences in boxplots.
- Unit 2 & Unit 9 (Regression & Slope Inference): Q5 was among the highest-scoring questions; however, students regularly forfeited points by interpreting regression slope deterministically rather than stating the predicted or estimated average change.
- Investigative Task (Q6): Successfully navigated by students who applied the Central Limit Theorem and sampling distribution standard error \(\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}\), but struggled when interpreting simulated empirical distributions of non-standard statistics (the sample range).
Examiner Pitfalls & Strategic Advice
- Definitive Language Trap: Always use non-deterministic phrasing. Never write that a hypothesis test "proves" a claim or that a slope "will increase by exactly \(x\)". Always state there is convincing statistical evidence or that the predicted value changes.
- Inference Procedure Identification: Carefully distinguish between the mean of differences (paired data, one-sample \(t\)-procedure) and the difference of two independent means (two-sample \(t\)-procedure).
- Direct Comparative Statements: When asked to compare distributions, avoid listing summaries in isolation. Explicitly use comparative conjunctions (e.g., "the median of cold streams is greater than the median of warm streams").