An original Thinka practice paper modelled on the structure and difficulty of the May 2025 AP AP Statistics paper. Not affiliated with or reproduced from AP.
Section Part A: Questions 1-5
Answer all five free-response questions. Show all work and clearly indicate methods used. Spend approximately 65 minutes on this section.
5 Question · 20 marks
Question 1 · Constructed Response
4 marks
A municipal environmental agency monitors daily particulate matter (\(\text{PM}_{2.5}\)) concentrations, measured in micrograms per cubic meter (\(\mu\text{g/m}^3\)), across different areas of a city. Independent random samples of 60 daily measurements were collected during the past year at two monitoring locations: Station North (near an industrial park) and Station South (near a suburban reserve).
The five-number summaries and outlier information for the daily \(\text{PM}_{2.5}\) concentrations at each station are shown in the table below:
A. Compare the distributions of daily \(\text{PM}_{2.5}\) concentrations for the sample from Station North and the sample from Station South.
B. For the sample of \(\text{PM}_{2.5}\) concentrations from Station North, would you expect the mean concentration to be greater than \(38\text{ }\mu\text{g/m}^3\), less than \(38\text{ }\mu\text{g/m}^3\), or equal to \(38\text{ }\mu\text{g/m}^3\)? Justify your answer.
C. Suppose the environmental agency combines the two samples into a single dataset of 120 daily observations.
i. What is the range of the combined dataset? Justify your answer.
ii. Explain why the median daily \(\text{PM}_{2.5}\) concentration of the combined dataset must be between \(22\text{ }\mu\text{g/m}^3\) and \(38\text{ }\mu\text{g/m}^3\).
Show answer & marking schemeHide answer & marking scheme
Worked solution
### Part A - Center: The median daily \(\text{PM}_{2.5}\) concentration for Station North (\(38\text{ }\mu\text{g/m}^3\)) is greater than the median daily concentration for Station South (\(22\text{ }\mu\text{g/m}^3\)). - Spread: The variability of daily \(\text{PM}_{2.5}\) concentrations is greater for Station North than for Station South. The range for Station North (\(85 - 10 = 75\text{ }\mu\text{g/m}^3\)) is greater than the range for Station South (\(44 - 8 = 36\text{ }\mu\text{g/m}^3\)), and the interquartile range (IQR) for Station North (\(46 - 22 = 24\text{ }\mu\text{g/m}^3\)) is greater than the IQR for Station South (\(30 - 14 = 16\text{ }\mu\text{g/m}^3\)). - Outliers / Shape: The distribution for Station North has an upper outlier at \(85\text{ }\mu\text{g/m}^3\) and is skewed to the right, whereas the distribution for Station South contains no outliers and is approximately symmetric. - Context: Stated clearly in terms of daily \(\text{PM}_{2.5}\) concentrations at Station North and Station South.
---
### Part B The mean daily \(\text{PM}_{2.5}\) concentration for Station North is expected to be **greater than \(38\text{ }\mu\text{g/m}^3\).
Justification:** The value \(38\text{ }\mu\text{g/m}^3\) is the median of the distribution. Because the distribution for Station North is right-skewed and has an extreme high outlier at \(85\text{ }\mu\text{g/m}^3\), the mean (which is not resistant to extreme values) will be pulled upward toward the higher values in the right tail, making the mean greater than the median.
---
### Part C i. $$\text{Range} = \text{Maximum} - \text{Minimum} = 85 - 8 = 77\text{ }\mu\text{g/m}^3$$ Justification: The overall maximum of the combined dataset is the maximum value observed across both stations (\(85\text{ }\mu\text{g/m}^3\) from Station North), and the overall minimum is the minimum value observed across both stations (\(8\text{ }\mu\text{g/m}^3\) from Station South).
ii. - In Station North, \(38\text{ }\mu\text{g/m}^3\) is the median, meaning at least \(30\) of its \(60\) values are \(\le 38\). In Station South, the maximum is \(44\) and the third quartile is \(30\), meaning far more than half (at least \(45\) values, and actually almost all \(60\) values since \(Q_3 = 30 < 38\)) are \(\le 38\). Thus, across the \(120\) combined observations, more than \(60\) observations (half of the data) are \(\le 38\), ensuring the combined median cannot exceed \(38\text{ }\mu\text{g/m}^3\). - Similarly, \(22\text{ }\mu\text{g/m}^3\) is the median of Station South (at least \(30\) values \(\ge 22\)) and is \(Q_1\) for Station North (at least \(45\) values \(\ge 22\)). Thus, in the combined dataset, at least \(30 + 45 = 75\) values (more than half of the \(120\) values) are \(\ge 22\), ensuring the combined median must be at least \(22\text{ }\mu\text{g/m}^3\). - Therefore, the median of the combined sample must lie between \(22\text{ }\mu\text{g/m}^3\) and \(38\text{ }\mu\text{g/m}^3\).
Marking scheme
### Scoring Guidelines
Each part is scored as Essentially Correct (E), Partially Correct (P), or Incorrect (I).
---
#### Part A is scored as: - Essentially Correct (E) if the response satisfies at least 3 of the following 4 components with appropriate context (referencing daily $\text{PM}_{2.5}$ concentrations and both stations): 1. Correctly compares the centers (medians) using comparative language (e.g., higher, greater). 2. Correctly compares the spread (IQR or range) using comparative language. 3. Identifies the outlier at $85\text{ }\mu\text{g/m}^3$ for Station North and notes that Station South has no outliers (or correctly compares shapes). 4. Includes full context (locations and units/variable). - Partially Correct (P) if the response satisfies 2 of the 4 components. - Incorrect (I) if the response satisfies fewer than 2 components.
---
#### Part B is scored as: - Essentially Correct (E) if the response satisfies all 3 components: 1. States that the mean is expected to be greater than $38\text{ }\mu\text{g/m}^3$. 2. Identifies $38\text{ }\mu\text{g/m}^3$ as the median of Station North. 3. Provides a valid justification based on the non-resistance of the mean to right-skewness and/or the high outlier at $85\text{ }\mu\text{g/m}^3$. - Partially Correct (P) if the response satisfies component 1 and either component 2 or component 3. - Incorrect (I) if the response does not meet the criteria for E or P.
---
#### Part C is scored as: - Essentially Correct (E) if both sub-parts (i) and (ii) are answered correctly: 1. (i) Calculates the range as $77\text{ }\mu\text{g/m}^3$ and justifies it using maximum = $85$ and minimum = $8$. 2. (ii) Correctly argues why the combined median must fall between $22$ and $38$ based on the sample sizes and the relative position of values/quartiles from both groups. - Partially Correct (P) if only one sub-part ((i) or (ii)) is fully correct with justification. - Incorrect (I) if neither sub-part is correctly answered and justified.
---
### Composite Score - 4 (Complete Response): All three parts (A, B, C) scored E. - 3 (Substantial Response): Two parts scored E and one part scored P. - 2 (Developing Response): Two parts scored E and one scored I; OR one part scored E and one or two scored P; OR all three parts scored P. - 1 (Minimal Response): One part scored E and two scored I; OR two parts scored P and one scored I.
Question 2 · Constructed Response
4 marks
A forestry manager wants to estimate the proportion of pine trees in a tree plantation that are infected with a needle fungus. The plantation is divided into a grid of 20 equal-sized sections arranged in 4 rows (labeled Row 1 through Row 4, from North to South) and 5 columns (labeled Column A through Column E, from West to East). Each section contains approximately 200 pine trees.
A busy interstate highway runs directly along the northern edge of the plantation, adjacent to Row 1. The manager believes that air pollution from the highway increases tree stress, making the proportion of infected trees highest in Row 1 and progressively lower in rows located farther south.
The manager is considering three different methods for selecting a sample of sections in order to inspect all pine trees within the chosen sections:
- Method I: Select only Section 1A (the northwest corner section, adjacent to the highway). Inspect every tree in Section 1A. - Method II: Randomly select one entire row from Rows 1, 2, 3, and 4. Inspect every tree in all 5 sections belonging to the selected row. - Method III: Randomly select one section from each of the 4 rows (Row 1, Row 2, Row 3, and Row 4). Inspect every tree in each of the 4 selected sections.
A. Explain whether Method I is an appropriate sampling method for the manager to use to estimate the proportion of infected pine trees in the entire plantation.
B. Suppose the manager uses Method II and randomly selects Row 1. If the manager’s belief about the effect of the highway is correct, determine whether this sample is likely to provide an overestimate or an underestimate of the true proportion of infected pine trees in the entire plantation. Justify your answer.
C. Describe how to implement Method III using a random number generator to select one section from each of the four rows.
Show answer & marking schemeHide answer & marking scheme
Worked solution
### Part A Method I is not an appropriate sampling method. It is a convenience sample that lacks random selection. Because Section 1A is adjacent to the interstate highway, where tree stress and fungal infection are believed to be highest, trees in this section are not representative of all trees in the plantation. Using only Section 1A is likely to result in substantial bias (an overestimate of the overall proportion of infected pine trees).
---
### Part B Selecting Row 1 using Method II is likely to provide an overestimate of the true proportion of infected pine trees in the entire plantation.
Justification: Row 1 is the row located closest to the busy interstate highway. According to the manager's belief, proximity to highway pollution increases fungal infection rates, meaning Row 1 has a higher proportion of infected trees than Rows 2, 3, and 4. Because every inspected tree comes exclusively from this most heavily infected row, the sample proportion will tend to exceed the overall plantation proportion.
---
### Part C To implement Method III using a random number generator:
1. Label the sections within each row: For each of the 4 rows (Row 1, Row 2, Row 3, Row 4), assign each of the 5 sections an integer from 1 to 5 (e.g., Column A = 1, Column B = 2, Column C = 3, Column D = 4, Column E = 5). 2. Generate random selections independently for each row: - For Row 1, use a random number generator to produce a single random integer from 1 to 5. Select the corresponding section in Row 1. - Repeat this procedure separately for Row 2 by generating a new random integer from 1 to 5 to select one section in Row 2. - Repeat separately for Row 3 and Row 4, generating one random integer from 1 to 5 for each row. 3. Inspect the selected sample: Inspect all 200 pine trees in each of the 4 randomly chosen sections (one per row) for fungal infection.
Marking scheme
### Scoring Guidelines
Each part is scored as Essentially Correct (E), Partially Correct (P), or Incorrect (I).
#### Part A - Essentially Correct (E): The response: 1. States that Method I is not appropriate; 2. Explains that the sample is non-random / a convenience sample OR not representative of the plantation; 3. References the context of the plantation and the variable (infected trees / fungal infection). - Partially Correct (P): The response satisfies two of the three components. - Incorrect (I): The response meets at most one component.
#### Part B - Essentially Correct (E): The response: 1. Correctly identifies that Row 1 will likely produce an overestimate; 2. Justifies this by noting that Row 1 is closest to the highway; 3. Explicitly connects the proximity to the highway with higher expected rates/proportions of fungal infection relative to the rest of the plantation. - Partially Correct (P): The response satisfies two of the three components (e.g., identifies overestimate and highway location but fails to clearly link higher disease rates to the rest of the field). - Incorrect (I): The response meets at most one component.
#### Part C - Essentially Correct (E): The response: 1. Clearly defines a numbering/labeling system for the sections within each stratum (row); 2. Describes a valid random number generator process ensuring all 5 sections within each row have an equal chance of selection; 3. Explicitly indicates that one section is selected independently from each of the 4 rows (resulting in a total sample of 4 sections). - Partially Correct (P): The response satisfies two of the three components. - Incorrect (I): The response satisfies at most one component.
---
### Holistic Score Conversion - 4 Points (Complete Response): All three parts E. - 3 Points (Substantial Response): Two parts E and one part P. - 2 Points (Developing Response): Two parts E and no part P, OR one part E and two parts P, OR three parts P. - 1 Point (Minimal Response): One part E and no parts P, OR two parts P.
Question 3 · free-response
4 marks
A boutique coffee company creates specialty sampler packs from a master inventory of 500 individually sealed single-origin coffee pods. The master inventory consists of four distinct varieties in the following quantities: 150 Ethiopian Yirgacheffe, 100 Colombian Supremo, 200 Guatemalan Antigua, and 50 Sumatran Mandheling. An automated packing dispenser selects pods at random with replacement from this master inventory to fill each promotional sampler pack.
A. i. Suppose one pod is selected at random from the master inventory. What is the probability that the pod is a Sumatran Mandheling pod? Show your work. ii. Suppose two pods are selected independently at random from the master inventory. What is the probability that both pods are Sumatran Mandheling pods? Show your work.
B. Each promotional sampler pack contains 15 randomly selected pods. A quality control coordinator is interested in the number of Sumatran Mandheling pods included in a typical promotional sampler pack. i. Define the random variable of interest to the coordinator, and state how the random variable is distributed, including its parameters. ii. What is the expected value of the random variable defined in part B (i)? Show your work.
C. Recall that each promotional sampler pack contains 15 pods chosen at random with replacement from the master inventory. i. Determine the probability that a randomly chosen promotional sampler pack contains 3 or more Sumatran Mandheling pods. Show your work. ii. A customer opens a promotional sampler pack and finds 3 Sumatran Mandheling pods. Does this observation provide strong evidence that the automated packing dispenser is not selecting pods at random from the master inventory? Justify your answer without performing a formal inference procedure.
Show answer & marking schemeHide answer & marking scheme
Worked solution
### Part A i. The probability that a single randomly selected pod is Sumatran Mandheling is: \[ P(\text{Sumatran}) = \frac{50}{500} = 0.10 \]
ii. Because the selections are independent (selected at random with replacement): \[ P(\text{Both Sumatran}) = P(\text{Sumatran}) \times P(\text{Sumatran}) = (0.10)(0.10) = 0.01 \]
---
### Part B i. Let the random variable \(X\) represent the number of Sumatran Mandheling pods in a promotional sampler pack of 15 pods. Because each pod selection is independent with a constant probability of success \(p = \frac{50}{500} = 0.10\), the random variable \(X\) follows a binomial distribution with parameters \(n = 15\) and \(p = 0.10\), denoted as \(X \sim \text{Binomial}(n = 15, p = 0.10)\).
ii. The expected value for the number of Sumatran Mandheling pods in a pack is: \[ E(X) = np = (15)(0.10) = 1.5\text{ pods} \]
---
### Part C i. The probability that 3 or more Sumatran Mandheling pods are in a sampler pack is: \[ P(X \ge 3) = 1 - P(X \le 2) = 1 - \sum_{k=0}^{2} \binom{15}{k} (0.10)^k (0.90)^{15-k} \] \[ P(X = 0) = \binom{15}{0}(0.10)^0(0.90)^{15} \approx 0.2059 \] \[ P(X = 1) = \binom{15}{1}(0.10)^1(0.90)^{14} \approx 0.3432 \] \[ P(X = 2) = \binom{15}{2}(0.10)^2(0.90)^{13} \approx 0.2669 \] \[ P(X \le 2) \approx 0.2059 + 0.3432 + 0.2669 = 0.8160 \] \[ P(X \ge 3) = 1 - 0.8160 = 0.1840 \text{ (or } 0.1841 \text{ using unrounded binomial values)} \]
ii. No, observing 3 Sumatran Mandheling pods does not provide strong evidence that the dispenser is not selecting pods at random. The probability of getting 3 or more Sumatran pods purely by chance is approximately \(0.184\) (or \(18.4\%\)). Because this probability is relatively high, obtaining 3 pods is not a rare occurrence and can reasonably be attributed to ordinary random variation.
Marking scheme
### Scoring Guidelines
Each of the three parts (A, B, C) is scored as Essentially Correct (E), Partially Correct (P), or Incorrect (I).
#### Part A is scored as follows: - Essentially Correct (E) if the response satisfies all 4 components: 1. Correctly calculates \(P = 0.10\) in part A (i). 2. Shows supporting work for part A (i). 3. Correctly calculates \(P = 0.01\) in part A (ii) consistent with part A (i). 4. Shows supporting work for part A (ii) demonstrating independence. - Partially Correct (P) if the response satisfies 2 or 3 of the 4 components. - Incorrect (I) if the response satisfies fewer than 2 components.
#### Part B is scored as follows: - Essentially Correct (E) if the response satisfies at least 3 of the 4 components: 1. Defines the random variable in context (number of Sumatran pods in a pack of 15 / in one pack). 2. Identifies the distribution as binomial. 3. States the parameters \(n = 15\) and \(p = 0.10\). 4. Correctly calculates the expected value \(E(X) = 1.5\) with supporting work. - Partially Correct (P) if the response satisfies 2 of the 4 components. - Incorrect (I) if the response satisfies fewer than 2 components.
#### Part C is scored as follows: - Essentially Correct (E) if the response satisfies all 4 components: 1. Correctly calculates the probability \(P(X \ge 3) \approx 0.184\) (or \(0.1841\)). 2. Shows supporting work (e.g., binomial formula, clearly labeled calculator syntax \(1 - \text{binomcdf}(n=15, p=0.1, x=2)\), or listing of probabilities). 3. Concludes that there is NOT strong evidence that the machine is not selecting randomly (or answers 'no'). 4. Provides justification linking the decision to the calculated probability being reasonably large / not rare / likely to happen by chance. - Partially Correct (P) if the response satisfies 2 or 3 of the 4 components. - Incorrect (I) if the response satisfies fewer than 2 components.
---
### Composite Score Conversion - 4 points (Complete Response): All three parts (A, B, C) scored E. - 3 points (Substantial Response): Two parts scored E and one part scored P. - 2 points (Developing Response): Two parts scored E and no part scored P; OR one part scored E and one/two parts scored P; OR all three parts scored P. - 1 point (Minimal Response): One part scored E and no parts scored P; OR no part scored E and two parts scored P. - 0 points: Response does not meet criteria for a score of 1.
Question 4 · free_response
4 marks
A national health survey reports that 35 percent of adults in a certain country get at least 8 hours of sleep per night. Dr. Ellis, a researcher in a large metropolitan city with more than 5,000 healthcare workers, believes that the proportion of adult healthcare workers in her city who get at least 8 hours of sleep per night is less than the national proportion of 0.35. To investigate her belief, Dr. Ellis selected a simple random sample of 150 adult healthcare workers from the city and found that 42 of them get at least 8 hours of sleep per night.
Is there convincing statistical evidence, at a significance level of \(\alpha = 0.05\), to support Dr. Ellis's belief? Justify your answer with an appropriate inference procedure.
Show answer & marking schemeHide answer & marking scheme
Worked solution
Step 1: State hypotheses and identify procedure
Let \(p\) represent the true proportion of all adult healthcare workers in the city who get at least 8 hours of sleep per night.
We test the hypotheses: \[H_0: p = 0.35\] \[H_a: p < 0.35\]
The appropriate procedure is a **one-sample \(z\)-test for a population proportion.
---
Step 2: Check conditions and compute test statistic and \(p\)-value**
- Random: A simple random sample of 150 healthcare workers from the city was taken. - 10% Condition: The sample size \(n = 150\) is less than 10% of all healthcare workers in the city since there are more than 5,000 healthcare workers (\(150 \le 0.10(5000) = 500\)). - Large Counts Condition: \[n p_0 = 150(0.35) = 52.5 \ge 10\] \[n(1 - p_0) = 150(1 - 0.35) = 150(0.65) = 97.5 \ge 10\] Because both expected counts are at least 10, the sampling distribution of \(\hat{p}\) is approximately normal.
**\(p\)-value:** \[p\text{-value} = P(Z \le -1.80) \approx 0.0359 \quad (\text{or } 0.0361 \text{ using unrounded } z = -1.7975)\]
---
Step 3: State conclusion in context
Because the \(p\)-value (\(\approx 0.036\)) is less than the significance level \(\alpha = 0.05\), we reject the null hypothesis \(H_0\).
There is convincing statistical evidence that the true proportion of adult healthcare workers in this city who get at least 8 hours of sleep per night is less than the national proportion of 0.35.
Marking scheme
This question is scored in three sections:
Section 1: Hypotheses and Identification of Procedure - Essentially correct (E) if the response satisfies all 4 components: 1. Identifies the inference procedure as a one-sample \(z\)-test for a population proportion (by name or formula). 2. States the null hypothesis as \(H_0: p = 0.35\). 3. States the alternative hypothesis as \(H_a: p < 0.35\). 4. Defines the parameter in context (the true proportion of adult healthcare workers in the city who get at least 8 hours of sleep per night). - Partially correct (P) if the response satisfies 3 of the 4 components. - Incorrect (I) if the response satisfies fewer than 3 components.
Section 2: Conditions and Calculations - Essentially correct (E) if the response satisfies all 4 components: 1. Checks the independence/random condition (mentions simple random sample and verifies the 10% condition: \(150 \le 0.10(5000)\)). 2. Verifies the Large Counts condition by calculating expected counts: \(np_0 = 150(0.35) = 52.5 \ge 10\) and \(n(1 - p_0) = 150(0.65) = 97.5 \ge 10\). 3. Calculates the correct \(z\)-statistic (\(z \approx -1.80\)). 4. Calculates a correct \(p\)-value consistent with the test statistic and one-sided alternative (\(p\text{-value} \approx 0.036\)). - Partially correct (P) if the response satisfies 2 or 3 of the 4 components. - Incorrect (I) if the response satisfies fewer than 2 components.
Section 3: Decision and Conclusion - Essentially correct (E) if the response satisfies both components: 1. Compares the \(p\)-value to \(\alpha = 0.05\) and makes the correct decision to reject \(H_0\). 2. States a correct conclusion in context in terms of the alternative hypothesis using non-definitive language. - Partially correct (P) if the response satisfies only 1 of the 2 components. - Incorrect (I) if neither component is satisfied.
Overall Question Score Distribution: - 4 points: Complete Response (EEE) - 3 points: Substantial Response (EEP) - 2 points: Developing Response (EEI, EPP, or PPP) - 1 point: Minimal Response (EII, PPI) - 0 points: Response that does not meet the criteria for 1 point.
Question 5 · Constructed Response
4 marks
According to a 2018 regional survey in Region K, the mean number of weekly volunteer hours contributed per active member of community garden cooperatives was 4.2 hours. Dylan, a researcher, believes that the mean number of weekly volunteer hours contributed per active member in Region K was different in 2024 than it was in 2018. To investigate his belief, Dylan obtained a large random sample of active members in Region K in 2024 and recorded the number of volunteer hours each member contributed during a typical week. The distribution of the number of weekly volunteer hours for the sampled members is summarized in the table below.
A. i. A member from the 2024 sample will be selected at random. What is the probability that the selected member volunteered more than 3 hours? Show your work. ii. What is the mean number of weekly volunteer hours for the sample of active members in 2024? Show your work.
B. Dylan will use a one-sample $t$-test for a population mean to test his belief. i. In the context of Dylan's investigation, state the null and alternative hypotheses. ii. Explain, in context, what a Type II error would be for Dylan's hypothesis test.
C. An independent researcher, Maya, suggests using a confidence interval to investigate whether the mean number of weekly volunteer hours in 2024 in Region K was different from 4.2 hours. Assume the conditions for inference have been met. Using Dylan's sample data, Maya calculated a one-sample 96 percent confidence interval to estimate the population mean as $(3.31, 3.73)$. Based on the confidence interval, what conclusion can be made for Dylan's hypothesis test in part B at $\alpha = 0.04$? Justify your answer.
Show answer & marking schemeHide answer & marking scheme
Worked solution
### Part A
i. Let $X$ represent the number of weekly volunteer hours for a randomly selected member from the 2024 sample. $$P(X > 3) = P(X = 4) + P(X = 5) + P(X = 6) = 0.22 + 0.14 + 0.06 = 0.42$$
ii. The sample mean number of weekly volunteer hours is calculated as: $$\bar{x} = E(X) = \sum x \cdot P(X = x)$$ $$\bar{x} = 1(0.08) + 2(0.15) + 3(0.35) + 4(0.22) + 5(0.14) + 6(0.06)$$ $$\bar{x} = 0.08 + 0.30 + 1.05 + 0.88 + 0.70 + 0.36 = 3.37\text{ hours}$$
---
### Part B
i. Let $\mu$ represent the population mean number of weekly volunteer hours contributed by active members of community garden cooperatives in Region K in 2024. - $H_0: \mu = 4.2$ - $H_a: \mu eq 4.2$
ii. A Type II error occurs when the researcher fails to reject a false null hypothesis. In this context, a Type II error would be concluding that the mean number of weekly volunteer hours for active garden members in 2024 is equal to 4.2 hours (or failing to find convincing evidence that it is different from 4.2 hours), when in fact the true population mean number of weekly volunteer hours in 2024 is different from 4.2 hours.
---
### Part C
Because the hypothesized value of $4.2$ is not included in the 96% confidence interval $(3.31, 3.73)$, the null hypothesis $H_0: \mu = 4.2$ should be rejected at the $\alpha = 1 - 0.96 = 0.04$ significance level. There is convincing statistical evidence that the true population mean number of weekly volunteer hours contributed by active community garden members in Region K in 2024 is different from 4.2 hours.
Marking scheme
Part A is scored as follows: - Essentially correct (E) if the response satisfies at least 3 of the following 4 components: 1. Correctly calculates the probability $0.42$ in part A (i). 2. Provides supporting work or equation for component 1. 3. Correctly calculates the sample mean $3.37$ in part A (ii). 4. Provides supporting work showing the sum of products for component 3. - Partially correct (P) if the response satisfies only 2 of the 4 components. - Incorrect (I) if the response does not meet the criteria for E or P.
Part B is scored as follows: - Essentially correct (E) if the response satisfies the following 4 components: 1. States the correct null hypothesis ($H_0: \mu = 4.2$ or in words) with the value 4.2. 2. States the correct two-sided alternative hypothesis ($H_a: \mu eq 4.2$ or in words). 3. Defines the parameter $\mu$ with context including reference to the population mean, variable (weekly volunteer hours), and sampling units/population (active members in Region K in 2024). 4. Correctly describes a Type II error in context, including the condition that the null hypothesis is actually false. - Partially correct (P) if the response satisfies only 3 of the 4 components. - Incorrect (I) if the response does not meet the criteria for E or P.
Part C is scored as follows: - Essentially correct (E) if the response satisfies both of the following components: 1. States a correct decision (reject $H_0$) and conclusion in context consistent with the two-sided alternative hypothesis using nondefinitive language. 2. Justifies the conclusion based on the value $4.2$ not being contained in the 96% confidence interval $(3.31, 3.73)$. - Partially correct (P) if the response satisfies only 1 of the 2 components. - Incorrect (I) if the response does not meet the criteria for E or P.
Final Score Scale: - 4 points: EEE - 3 points: EEP - 2 points: EEI, EPP, or PPP - 1 point: EPI, PPI, or EII - 0 points: PII or III
Ready to test yourself?
Turn these notes into exam-style practice. Get unlimited AI questions on this topic with instant marking and explanations.
Answer Question 6. This question requires extending course concepts to an unfamiliar context. Spend approximately 25 minutes on this question.
1 Question · 4 marks
Question 1 · Extended Investigative Task
4 marks
6. Dr. Liang, an environmental scientist, conducted a study to investigate the difference in microplastic contamination between two freshwater lakes, Lake North and Lake South. Water samples of 1 liter each were collected from 40 randomly selected locations across Lake North and 40 randomly selected locations across Lake South at identical depths and times. For each water sample, the concentration of microplastics (in particles per liter) was recorded. Dr. Liang is interested in comparing the mean microplastic concentration between the two lakes. Table 1 shows the summary statistics from the study.
Dr. Liang confirmed that conditions for inference were met and conducted a two-sample $t$-test for the difference in two population means. Let $\mu_N$ represent the mean microplastic concentration (particles/L) for all possible sample locations in Lake North, and let $\mu_S$ represent the mean microplastic concentration (particles/L) for all possible sample locations in Lake South.
A. The $p$-value for Dr. Liang's hypothesis test was $0.0006$. State an appropriate conclusion, at the $1$ percent significance level, for Dr. Liang's test in the context of the study. Justify your answer.
B. Explain why it was appropriate for Dr. Liang to conduct a two-sample $t$-test for the difference in two population means instead of a paired $t$-test for a population mean difference.
C. In environmental monitoring, researchers often evaluate practical disparity in addition to statistical significance. One metric used to assess the magnitude of practical difference between two independent environmental distributions is the **Relative Disparity Index ($RDI$)**.
The $RDI$ is calculated using: $$RDI = \frac{|\bar{x}_1 - \bar{x}_2|}{s_{\text{comb}}}$$ where $\bar{x}_1$ and $\bar{x}_2$ are the sample means of the two groups, and $s_{\text{comb}}$ is the combined root variance defined when sample sizes are equal by: $$s_{\text{comb}} = \sqrt{s_1^2 + s_2^2}$$ where $s_1$ and $s_2$ represent the sample standard deviations for the two groups.
i. Calculate the $RDI$ for Dr. Liang's study. Show your work.
ii. Higher values of $RDI$ indicate greater practical disparity between ecosystems. Table 2 provides general guidelines for interpreting the $RDI$.
Based on your answer to part C (i) and the information in Table 2, describe the practical disparity in microplastic contamination between Lake North and Lake South, in context.
D. Suppose a second study conducted on two different lakes yielded the same sample means ($\bar{x}_1 = 28.4$ and $\bar{x}_2 = 23.6$) and the same sample sizes ($n_1 = n_2 = 40$), but the standard deviation for both lakes was greater than $9.00$.
i. Would the $RDI$ in this new situation be smaller than, larger than, or the same as the $RDI$ calculated in part C (i)? Explain your answer.
ii. Does the $RDI$ described in part D (i) indicate that the observed difference in means in the new situation has more practical disparity, less practical disparity, or the same practical disparity compared to what was determined in part C (ii)? Explain your answer.
Show answer & marking schemeHide answer & marking scheme
Worked solution
Part A: Because the $p$-value of $0.0006$ is less than the significance level $\alpha = 0.01$, we reject the null hypothesis $H_0$. There is convincing statistical evidence that the true mean microplastic concentration for all sample locations in Lake North ($\mu_N$) is different from the true mean microplastic concentration for all sample locations in Lake South ($\mu_S$).
Part B: A two-sample $t$-test is appropriate because the data consist of two independent random samples of water locations taken from two separate lakes. There is no natural pairing or one-to-one matching between individual water samples from Lake North and Lake South (e.g., they are not paired by identical spatial coordinates, nor is each sample measured before and after a treatment).
ii. Based on Table 2, an $RDI$ value of $0.565$ falls in the interval $0.30 < RDI < 0.70$, indicating that there is moderate environmental disparity in the microplastic contamination levels between Lake North and Lake South.
Part D: i. The $RDI$ in this new situation would be smaller than the $RDI$ calculated in part C (i). Reason: The difference between the sample means $|\bar{x}_1 - \bar{x}_2| = |28.4 - 23.6| = 4.8$ remains unchanged in the numerator. However, since the standard deviations $s_1$ and $s_2$ are both greater than $9.00$, the denominator $s_{\text{comb}} = \sqrt{s_1^2 + s_2^2} > \sqrt{9.00^2 + 9.00^2} = \sqrt{162} \approx 12.73$, which is larger than $8.49$. Increasing the denominator while keeping the numerator constant results in a smaller quotient for $RDI$.
ii. It indicates less practical disparity. Reason: A lower $RDI$ value indicates that the observed difference between the means is smaller relative to the natural variability (noise) of microplastic levels within the lakes, representing less meaningful real-world separation between the two ecosystem distributions.
Marking scheme
Part A is scored as follows: - Essentially correct (E) if the response: 1. Correctly compares the $p$-value ($0.0006$) to $\alpha = 0.01$ and states a correct decision regarding $H_0$ (reject $H_0$). 2. States a correct conclusion in context (microplastic concentration, Lake North vs. Lake South) in terms of the alternative hypothesis using non-definitive language. - Partially correct (P) if the response satisfies only 1 of the 2 components. - Incorrect (I) if neither component is satisfied.
Part B is scored as follows: - Essentially correct (E) if the response: 1. States that the two groups/samples are independent (or not paired). 2. Justifies independence/non-pairing in context (e.g., separate random sampling locations in two distinct lakes rather than matched pairs or repeated measures on the same units). - Partially correct (P) if only 1 of the 2 components is satisfied. - Incorrect (I) if neither component is satisfied.
Part C is scored as follows: - Essentially correct (E) if the response: 1. Correctly calculates $RDI \approx 0.565$ (or values rounding between $0.56$ and $0.57$) with supporting work shown in subpart (i). 2. Correctly identifies and describes the practical importance as 'moderate environmental disparity' in context of microplastic concentrations in the two lakes in subpart (ii). - Partially correct (P) if only 1 of the 2 components is satisfied. - Incorrect (I) if neither component is satisfied.
Part D is scored as follows: - Essentially correct (E) if the response: 1. In (i), correctly identifies that $RDI$ would be smaller AND provides a valid justification based on an increased denominator ($s_{\text{comb}}$) with a constant numerator. 2. In (ii), correctly identifies that this represents less practical disparity AND justifies this based on the properties of $RDI$ / Table 2. - Partially correct (P) if only 1 of the 2 components is satisfied. - Incorrect (I) if neither component is satisfied.
Overall Question Score (0–4): - 4 points: 4 E's (or 3 E's and 1 P with strong holistic communication) - 3 points: 3 E's, or 2 E's and 2 P's - 2 points: 2 E's, or 1 E and 2 P's, or 4 P's - 1 point: 1 E and 0 P's, or 2–3 P's - 0 points: 0 E's and 0–1 P
Wondering how well you actually know this?
thinka is an AI practice app for DSE students: unlimited questions, instant auto-marking, and detailed step-by-step solutions. 100,000+ students use it to confirm they actually know it, not just think they do.