An original Thinka practice paper modelled on the structure and difficulty of the Jun 2025 CCEA GCSE Statistics 2260 paper. Not affiliated with or reproduced from CCEA.
Section Unit 1: Core Processing, Probability, Quality Control & Normal Modeling
Answer all ten questions. Write answers in the spaces provided. Calculator allowed. Formula sheet on page 2.
10 Question · 100 marks
Question 1 · Outlier detection, stem & leaf diagram, median and range
8 marks
The times, in minutes, taken by 15 students to complete a puzzle were recorded as:
(a) Draw an ordered stem and leaf diagram to represent this data, using stems of 2, 3, 4, 5 and 6, and include a key. [3] (b) Find the median and the range of the data. [2] (c) Calculate the interquartile range (IQR) of the data. [2] (d) Using the rule that an upper outlier is a value greater than UQ + 1.5×IQR, determine whether the value 62 is an outlier. [1]
Show answer & marking schemeHide answer & marking scheme
(d) Upper outlier bound \( = UQ+1.5\times IQR = 43+1.5(12) = 43+18 = 61 \). Since 62 > 61, the value 62 IS classified as an outlier.
Final answer: median = 35, range = 39, IQR = 12, and 62 is an outlier.
Marking scheme
(a) M1 for stems 2-6 and leaves correctly sorted in ascending order within each row; A1 for all 15 leaves correctly placed; A1 for a correct key. (b) A1 median = 35; A1 range = 39. (c) M1 correct identification of LQ (=31) and UQ (=43) by splitting the ordered data into halves; A1 IQR = 12. (d) MA1 correct calculation of the upper bound (61) AND correct conclusion that 62 is an outlier (both required for the mark).
A student researching healthy eating habits designs the survey question: 'Don't you agree that everyone should eat at least five portions of fruit and vegetables a day?'
(a) Explain why this is a leading question and why this could lead to bias in the results. [2] (b) Rewrite the question as a closed question that would avoid this bias. [1]
A box plot for the number of portions of fruit and vegetables eaten per day by a sample of adults shows the following summary: minimum = 1, lower quartile = 3, median = 4, upper quartile = 6, maximum = 9.
(c) State the interquartile range and the range shown by this box plot. [2]
A compound percentage bar chart shows the breakdown of survey responses to 'How satisfied are you with local healthy eating facilities?' as follows:
Year Very satisfied Satisfied Unsatisfied 2023 30% 45% 25% 2024 40% 40% 20%
(d) State the percentage of respondents who were 'Unsatisfied' in 2024. [1] (e) Calculate the percentage point change in the proportion of respondents who were 'Very satisfied' between 2023 and 2024. [3]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) The question is leading because its phrasing ('Don't you agree that...') suggests the 'correct'/expected answer is 'yes', pressuring respondents towards agreeing rather than giving their genuine, independent opinion; this would bias the results towards showing artificially strong support for eating five portions a day.
(b) For example: 'How many portions of fruit and vegetables do you think an adult should eat per day? (0 / 1-2 / 3-4 / 5 or more)' — a neutrally-worded closed question offering fixed response options without suggesting a particular answer.
(a) A1 correctly identifies the question suggests/pressures towards a particular ('yes') answer; A1 correctly links this to biased/unrepresentative results. (b) A1 valid closed question with neutral wording and fixed response categories. (c) A1 IQR = 3; A1 range = 8. (d) A1 20%. (e) M1 correct method (2024 value minus 2023 value); A1 correct value 10; A1 correctly states this as an increase/'percentage point' change (not a % change).
Question 3 · Biased coin tree diagram, outcome listing, combined probability and expected frequency
11 marks
A biased coin has probability 0.6 of landing on heads (H) and probability 0.4 of landing on tails (T) on each toss. The coin is tossed twice.
(a) Complete a probability tree diagram for the two tosses, labelling all branch probabilities. [3] (b) List all four possible outcomes for the two tosses, and state the probability of each. [2] (c) Calculate the probability of obtaining exactly one head in the two tosses. [3] (d) Calculate the probability of obtaining at least one head in the two tosses. [1] (e) The coin is tossed twice, 200 times in total (i.e. 200 repetitions of the two-toss experiment). Calculate the expected number of repetitions in which exactly one head is obtained. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) First toss: P(H) = 0.6, P(T) = 0.4. From each of these, second toss: P(H) = 0.6, P(T) = 0.4 again (the coin has no memory), giving four branch pairs: H then H (0.6, 0.6); H then T (0.6, 0.4); T then H (0.4, 0.6); T then T (0.4, 0.4).
(a) Describe the trend in the price shown by this time series data. [2] (b) Using 2019 as the base year (index = 100), calculate the price index for each of the years 2020 to 2023. Give your answers to 2 decimal places. [4] (c) Interpret, in context, the index number you calculated for 2022. [2] (d) State one advantage of using index numbers, rather than the original prices, to compare price changes over time. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) The price rose every year over the five-year period, from £40 in 2019 to £54 in 2023 — the price shows a steady, continuous upward (increasing) trend.
(b) Index \( = \dfrac{\text{price in year}}{\text{price in base year}}\times100 \). 2020: \( \dfrac{42}{40}\times100=105.00 \). 2021: \( \dfrac{45}{40}\times100=112.50 \). 2022: \( \dfrac{50}{40}\times100=125.00 \). 2023: \( \dfrac{54}{40}\times100=135.00 \).
(c) An index of 125.00 for 2022 means that the price of the item in 2022 was 25% higher than its price in the base year, 2019.
(d) Index numbers allow price changes to be compared easily as percentages relative to a fixed base year, making it simple to see and compare the SIZE of relative changes over time (or between different items with very different original prices) without needing to work with the original, differently-scaled values each time.
Final answer: (b) 2020=105.00, 2021=112.50, 2022=125.00, 2023=135.00.
Marking scheme
(a) A1 correctly describes a continuous/steady increase; A1 correctly quotes the overall change from £40 to £54 (or equivalent supporting detail). (b) M1 correct formula/method shown; A1 2020 and 2021 correct; A1 2022 correct; A1 2023 correct (all four to 2 d.p.). (c) MA1 correct interpretation: price 25% higher than in 2019 (both the percentage AND the correct reference to the base year required). (d) A1 basic advantage stated; A1 fully-developed explanation (e.g. easy comparison of relative change/percentage terms, or comparability across different items/scales).
Question 5 · Stepped cumulative frequency, data classification, percentiles/deciles
11 marks
The number of pets owned by each of 40 households was recorded:
Number of pets 0 1 2 3 4 5 Frequency 8 12 10 6 3 1
(a) State whether the number of pets owned is discrete or continuous data, and explain why a stepped (rather than smooth) cumulative frequency diagram is appropriate for this data. [2] (b) Complete a cumulative frequency table for this data. [2] (c) Use your table to find the median number of pets owned. [2] (d) Find the 3rd decile and the 90th percentile of the data. [3] (e) Interpret, in context, the value you found for the 90th percentile. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) The number of pets owned is discrete data, since it can only take whole-number (integer) values (a household cannot own, for example, 2.5 pets). A stepped cumulative frequency diagram is appropriate because the cumulative frequency only changes (jumps) at each specific discrete value, remaining constant in between, rather than increasing smoothly and continuously as it would for grouped continuous data.
(c) With n=40, the median is the average of the 20th and 21st values. From the cumulative frequency table, the 20th value is the last value in the 'up to 1 pet' group (cumulative frequency reaches exactly 20 at 1 pet), so the 20th value = 1; the 21st value falls in the next group (2 pets), so the 21st value = 2. Median \( = \dfrac{1+2}{2} = 1.5 \) pets.
(d) 3rd decile position \( = 0.3\times40 = 12 \)th value; this falls within the '1 pet' group (cumulative frequency reaches 20 at 1 pet, and the 12th value lies between the 9th and 20th values), so the 3rd decile = 1. 90th percentile position \( = 0.9\times40 = 36 \)th value; this falls exactly at the boundary of the '3 pets' group (cumulative frequency reaches 36 at 3 pets), so the 90th percentile = 3.
(e) A 90th percentile of 3 means that 90% of the households in the sample own 3 pets or fewer (and correspondingly, only 10% of households own more than 3 pets).
Final answer: (c) median = 1.5; (d) 3rd decile = 1, 90th percentile = 3.
Marking scheme
(a) A1 correctly identifies discrete; A1 correct reasoning linking discreteness to the stepped (rather than smooth curve) shape. (b) A1 correct cumulative frequencies for pets 0-2 (8, 20, 30); A1 correct cumulative frequencies for pets 3-5 (36, 39, 40). (c) M1 correct identification of the 20th and 21st values using the table; A1 median = 1.5. (d) M1 correct position calculation for the 3rd decile (12th value) and correct value = 1; A1 correct position calculation for the 90th percentile (36th value) and correct value = 3 (M1 A1, or MA1 for each if combined). (e) MA1 correct interpretation referencing 90% of households owning 3 or fewer pets.
Question 6 · Bivariate scatter diagram, correlation interpretation, line of best fit equation and gradient context
10 marks
The number of hours studied (x) and the examination score (y, out of 100) achieved by 8 students are shown below.
You are given: \( S_{xx}=42.0 \), \( S_{yy}=1565.9 \), \( S_{xy}=255.5 \).
(a) Plot a scatter diagram for this data and describe the type of correlation shown. [3] (b) Calculate the product moment correlation coefficient, giving your answer to 3 decimal places, and interpret its value in context. [3] (c) Given that the equation of the line of best fit is of the form \( y=a+bx \), calculate the values of a and b, using the fact that the line passes through the mean point \( (\bar{x},\bar{y}) \). [3] (d) Interpret, in context, the value of b (the gradient) in your equation. [1]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) Points plotted accurately, with hours studied on the x-axis and exam score on the y-axis; as x increases, y also clearly tends to increase, showing strong positive correlation.
(b) \( r=\dfrac{S_{xy}}{\sqrt{S_{xx}S_{yy}}}=\dfrac{255.5}{\sqrt{42.0\times1565.9}}=\dfrac{255.5}{\sqrt{65\,767.8}}=\dfrac{255.5}{256.45}=0.996 \) (3 d.p.). Since r is very close to +1, this indicates a very strong positive linear correlation between hours studied and exam score.
(d) The gradient, b ≈ 6.08, means that, on average, each additional hour of studying is associated with an increase of approximately 6.08 marks in the examination score.
Final answer: \( r=0.996 \); \( y=28.17+6.08x \).
Marking scheme
(a) A1 correctly labelled axes and scale; A1 all 8 points plotted accurately; A1 correct description (strong positive correlation). (b) M1 correct substitution into the PMCC formula; A1 \( r=0.996 \); A1 correct interpretation (very strong positive linear correlation). (c) M1 correct method for b (\( S_{xy}/S_{xx} \)); A1 correct \( \bar{x} \), \( \bar{y} \) and correctly-derived value of a; A1 fully correct equation \( y=28.17+6.08x \) (accept values rounded to 2-3 d.p.). (d) A1 correct contextual interpretation of the gradient (marks gained per extra hour studied).
Question 7 · Statistical process control chart, target/warning/action lines, out-of-control sample evaluation
11 marks
A factory monitors the mean weight, in grams, of samples of 5 cereal boxes taken every hour from its production line. The target weight is 50 g. Warning lines are set at 45 g and 55 g, and action lines are set at 40 g and 60 g.
The sample means for 10 consecutive hourly samples were:
(a) Draw and label a control chart showing the target line, the two warning lines and the two action lines, and plot the 10 sample means. [3] (b) Identify any sample(s) that fall between a warning line and an action line, and any sample(s) that fall beyond an action line. [3] (c) State what action, if any, the factory should take following sample 7 and following sample 9, and explain your reasoning. [3] (d) Explain why quality control charts, using warning and action lines, are useful in a production process such as this. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) A control chart with weight (g) on the vertical axis and sample number on the horizontal axis, with horizontal lines drawn at: target = 50 g (centre line), warning lines at 45 g and 55 g, and action lines at 40 g and 60 g; all 10 sample means plotted as points at their correct sample number and value, usually joined by straight line segments.
(b) Sample 7 (58 g) lies between the upper warning line (55 g) and the upper action line (60 g) — a warning breach. Sample 9 (62 g) lies beyond the upper action line (60 g) — an action breach. All other samples (1-6, 8, 10) lie between the two warning lines and so are within normal, expected variation.
(c) Following sample 7 (a single warning breach), the process should continue to be monitored closely, but no immediate corrective action is required from a single warning-line breach alone (a warning signals the process may be starting to drift and should be watched, but is not yet conclusive evidence of a problem). Following sample 9 (an action breach, 62 g, beyond the action line), the factory should stop the production process and investigate immediately, since this indicates the process mean has very likely shifted out of control and boxes may not be meeting their target weight specification.
(d) Warning and action lines allow production staff to distinguish normal, expected random variation in sample weights (falling within the warning lines) from signs that the process may be drifting out of control (warning breach) or has definitely gone out of control (action breach), enabling problems to be identified and corrected promptly — helping to maintain consistent product quality (e.g. ensuring customers reliably receive the advertised weight of product) while avoiding unnecessary process stoppages for normal, minor fluctuations.
Final answer: Sample 7 triggers a warning; Sample 9 triggers an action (process should be stopped and investigated).
Marking scheme
(a) A1 correctly labelled axes; A1 all five reference lines (target, 2 warning, 2 action) correctly drawn and labelled; A1 all 10 points plotted accurately. (b) A1 correctly identifies sample 7 as a warning breach; A1 correctly identifies sample 9 as an action breach; A1 correctly confirms all other samples are within the warning lines. (c) A1 correct action following sample 7 (continue/monitor, no immediate stoppage) with valid reasoning; A1 correct action following sample 9 (stop and investigate the process); A1 valid, clearly-expressed reasoning for the sample 9 decision. (d) A1 basic explanation of the purpose of warning/action lines; A1 fully-developed explanation referencing distinguishing normal variation from a genuine process problem and maintaining product quality/consistency.
Question 8 · Grouped frequency distribution mean estimate, scaling effects on mean and standard deviation
9 marks
The ages, in completed years, of 35 people attending a fitness class are grouped below.
Age (years) 0-10 10-20 20-30 30-40 40-50 Midpoint 5 15 25 35 45 Frequency 4 9 12 7 3
(a) Calculate an estimate of the mean age of people attending the class. [3] (b) Explain why your answer to (a) is only an ESTIMATE of the mean, rather than the exact mean. [2] (c) Every person's age is increased by 3 years (one year for each of the next three birthdays, recorded consistently for the whole group). State the effect this would have on the mean and on the standard deviation of the ages, giving reasons. [2] (d) Suppose, instead, every person's age (in years) were converted to age in months (i.e. each value multiplied by 12). State the effect this would have on the mean and on the standard deviation of the ages, giving reasons. [2]
Show answer & marking schemeHide answer & marking scheme
(b) This is only an estimate because the exact individual ages within each class interval are not known; the calculation assumes every person in a class interval has an age exactly equal to that interval's midpoint, which is unlikely to be true for every individual (their actual ages are simply spread somewhere within each interval).
(c) Adding a constant (3 years) to every value increases the mean by exactly that same constant (new mean = old mean + 3), because the whole distribution shifts up by 3 years; however, the standard deviation is UNCHANGED, because adding a constant to every value does not change how spread out the values are relative to one another (the shape and spread of the distribution stays the same, it is simply shifted).
(d) Multiplying every value by a constant (12, to convert years to months) multiplies the mean by that same constant (new mean = old mean × 12); the standard deviation is also multiplied by that same constant (new SD = old SD × 12), because scaling every value also scales the spread of the data by the same factor.
Final answer: (a) mean estimate = 23.86 years.
Marking scheme
(a) M1 correct method (Σfx/Σf) with correct Σfx = 835 shown; A1 mean = 23.86 (or 23 6/7) years. (b) A1 correctly identifies individual ages within a class are unknown; A1 correctly explains the midpoint assumption underlying the estimate. (c) A1 correctly states mean increases by 3 (to new mean); A1 correctly states SD is unchanged, with valid reasoning. (d) A1 correctly states mean is multiplied by 12; A1 correctly states SD is also multiplied by 12, with valid reasoning.
(a) Explain why a 4-point moving average is appropriate for this quarterly data. [1] (b) Calculate the first two (uncentred) 4-point moving averages. [2] (c) By taking the mean of successive pairs of 4-point moving averages, calculate the centred 4-point moving average values, which will align with Q3-22, Q4-22, Q1-23 and Q2-23. [4] (d) Plot the original data and the trend line (centred moving averages) on the same axes, and use your trend line to comment on the underlying trend in sales. [2] (e) Explain one limitation of using this trend line to forecast sales for Q3-23 and beyond, before it was known. [1]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) A 4-point moving average is appropriate because there are 4 quarters in each year, so averaging every 4 consecutive values smooths out the seasonal variation between quarters, revealing the underlying trend.
(b) First 4-point MA (Q1-22 to Q4-22): \( \dfrac{40+55+70+50}{4}=53.75 \). Second 4-point MA (Q2-22 to Q1-23): \( \dfrac{55+70+50+45}{4}=55.00 \).
(c) Continuing, the uncentred 4-point MAs are: 53.75, 55.00, 56.25, 57.50, 58.75. Since a 4-point moving average falls between two quarters (an even number of points), pairs of consecutive uncentred MAs are averaged ('centring') to align each value with an actual quarter: \( \frac{53.75+55.00}{2}=54.375 \) (Q3-22); \( \frac{55.00+56.25}{2}=55.625 \) (Q4-22); \( \frac{56.25+57.50}{2}=56.875 \) (Q1-23); \( \frac{57.50+58.75}{2}=58.125 \) (Q2-23).
(d) The trend line values (54.375, 55.625, 56.875, 58.125) increase steadily from Q3-22 to Q2-23, indicating that, once seasonal quarterly fluctuations are smoothed out, underlying sales are showing a steady upward (increasing) trend over this period.
(e) Forecasting beyond the available data (extrapolation) assumes that the existing trend and seasonal pattern will continue unchanged into the future, which may not be true — unforeseen factors (e.g. a new competitor, economic downturn, change in consumer tastes) could cause actual future sales to deviate significantly from a simple extrapolation of the trend line.
Marking scheme
(a) A1 correct explanation referencing 4 quarters per year/smoothing seasonal variation. (b) A1 first MA = 53.75; A1 second MA = 55.00. (c) M1 correct method (averaging successive uncentred MA pairs); A1 Q3-22 = 54.375; A1 Q4-22 = 55.625; A1 Q1-23 = 56.875 and Q2-23 = 58.125 (both required for this final mark). (d) A1 correctly plots/describes both series; A1 correct comment identifying a steady increasing underlying trend. (e) A1 valid limitation of extrapolation (assumes trend/pattern continues; unforeseen factors could change the actual outcome).
Question 10 · Normal distribution sketching, standardized z-score outlier verification, empirical rule probabilities
11 marks
The lengths of a large batch of manufactured bolts are modelled by a normal distribution with mean \( \mu=100\ \text{mm} \) and standard deviation \( \sigma=15\ \text{mm} \).
(a) Sketch the shape of this normal distribution, marking the mean and indicating its symmetry. [2] (b) Calculate the standardised (z) score for a bolt of length 130 mm. [2] (c) A bolt is considered a manufacturing anomaly if its z-score is more than 3 in magnitude. Determine whether a bolt of length 130 mm would be classed as an anomaly, giving a reason. [2] (d) Using the fact that approximately 95% of values in a normal distribution lie within two standard deviations of the mean, estimate the probability that a randomly selected bolt has a length greater than 130 mm. [3] (e) Using the fact that 68% of values lie within one standard deviation of the mean, estimate the percentage of bolts with lengths between 85 mm and 115 mm. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) A symmetric, bell-shaped curve, centred on and symmetric about the mean, 100 mm, with the curve tailing off (approaching but never quite reaching the horizontal axis) on either side.
(c) Since \( |z|=2.0 \), which is less than 3, a bolt of length 130 mm would NOT be classed as an anomaly (its length, while somewhat unusual, is not extreme enough to be flagged, since values more than 3 standard deviations from the mean are considered very unusual).
(d) Approximately 95% of values lie within two standard deviations of the mean, i.e. between \( 100-2(15)=70 \) mm and \( 100+2(15)=130 \) mm. This leaves approximately \( 100\%-95\%=5\% \) of values outside this range, split (by symmetry) equally between the two tails: \( 5\%\div2=2.5\% \) of bolts have length greater than 130 mm.
(e) 85 mm and 115 mm are exactly one standard deviation below and above the mean respectively (\( 100-15=85 \), \( 100+15=115 \)); since 68% of values lie within one standard deviation of the mean, approximately 68% of bolts have lengths between 85 mm and 115 mm.
Final answer: (b) z=2.0; (c) not an anomaly; (d) 2.5%; (e) 68%.
Marking scheme
(a) A1 correct symmetric bell shape; A1 mean correctly marked at the centre/axis of symmetry. (b) M1 correct substitution into \( z=(x-\mu)/\sigma \); A1 z = 2.0. (c) A1 correctly states not an anomaly (ECF from (b)); A1 valid reasoning (|z|<3). (d) M1 correctly identifies 130 mm as 2 SD above the mean; M1 correct method for the tail probability (100-95, halved); A1 2.5%. (e) A1 correctly identifies 85 and 115 as one SD below/above the mean; A1 correct value, 68%.
Ready to test yourself?
Turn these notes into exam-style practice. Get unlimited AI questions on this topic with instant marking and explanations.
Section Unit 2: Advanced Data Analysis, Probability Models & Risk
Answer all eleven questions. Write answers in the spaces provided. Calculator allowed. Formula sheet on page 2.
11 Question · 100 marks
Question 1 · Comparative line graph analysis across demographic groups
4 marks
A line graph shows the average weekly screen time (hours) of male and female teenagers over four years:
Year 2020 2021 2022 2023 Male (hrs) 18 20 23 24 Female (hrs) 16 19 22 25
(a) Describe the trend in screen time for each group between 2020 and 2023. [2] (b) Compare the screen time of male and female teenagers over this period, identifying any point at which the relationship between the two groups changes. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) Both male and female average screen time increased steadily every year from 2020 to 2023 (male: 18 to 24 hours; female: 16 to 25 hours).
(b) In 2020, 2021 and 2022, male teenagers had higher average screen time than female teenagers (e.g. 18 vs 16 hours in 2020); however, female screen time increased at a faster rate each year, and by 2023 female screen time (25 hours) had overtaken male screen time (24 hours) — this is the point at which the relationship between the two groups changes.
Marking scheme
(a) A1 correctly describes both groups increasing steadily; A1 correct supporting start/end values for at least one group. (b) A1 correctly identifies male initially higher than female; A1 correctly identifies female overtakes male by 2023.
Question 2 · Critique of misleading charts and scale misrepresentations
3 marks
A company presents a bar chart of its annual profit, in £ millions, for the last three years: Year 1 = £9.8m, Year 2 = £10.0m, Year 3 = £10.4m. On the bar chart, the vertical axis starts at £9.5m (rather than £0m) and is not clearly labelled with this starting value, and the three bars are drawn using different widths.
List three problems with this bar chart that could mislead a reader.
Show answer & marking schemeHide answer & marking scheme
Worked solution
1. The vertical axis is truncated (does not start at zero), which visually exaggerates the size of the differences between the three years' profits, making relatively small actual changes look much larger than they really are. 2. The starting value of the truncated axis is not clearly labelled, so a reader glancing at the chart may not even realise the axis has been truncated, increasing the risk of being misled. 3. The bars are drawn with different widths, which is inappropriate for a bar chart (widths should be equal) and could visually suggest a difference in importance or magnitude between the years beyond what the height (value) alone represents.
Marking scheme
A1 each for any three distinct, correctly identified problems (truncated/non-zero axis; axis start value not clearly labelled; unequal bar widths); max [3]. All other valid problems (e.g. no scale markings at all, misleading 3D effect if described) credited.
Question 3 · Statistical enquiry cycle planning, hypothesis generation and data sourcing
8 marks
A student wants to investigate whether students who cycle to school have better cardiovascular fitness than students who are driven to school.
(a) Write a suitable hypothesis for this investigation. [2] (b) State whether the data the student needs to collect would be primary or secondary data, and describe one appropriate method of collecting it. [2] (c) Suggest an appropriate sampling method the student could use to select participants from their school, and explain why it is appropriate. [2] (d) Identify one variable, other than mode of travel to school, that the student should try to control or record, to avoid it confounding the results. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) A suitable hypothesis: 'Students who cycle to school have better cardiovascular fitness (e.g. a lower resting heart rate, or a better result in a standard fitness test) than students who are driven to school.'
(b) This would be primary data, since the student would need to collect new, original data specifically for this investigation (it does not already exist in a usable, pre-collected form). An appropriate method would be to carry out a standard fitness test (e.g. a timed run, or measuring resting heart rate/recovery rate after exercise) on a sample of students, alongside recording their usual mode of travel to school.
(c) A stratified sample, selecting proportionally from each year group (or from 'cyclists' and 'driven' groups identified in advance), would be appropriate, because it would ensure the sample fairly represents students across different year groups/ages, reducing the risk that age-related fitness differences bias the comparison between the two travel-mode groups.
(d) A relevant confounding variable to control or record would be the amount of other physical exercise/sport each student does outside of travelling to school (e.g. participation in sports clubs), since a student's overall fitness level could be affected by this rather than (or as well as) their mode of travel to school, potentially confounding the results if not accounted for. (Other valid answers: age, gender, distance travelled, general health conditions.)
Marking scheme
(a) A1 valid hypothesis relating mode of travel to fitness; A1 hypothesis is specific/measurable (references a particular fitness measure or clear comparison). (b) A1 correctly identifies primary data; A1 valid, specific data collection method described. (c) A1 valid sampling method named (e.g. stratified, systematic, random); A1 correct reasoning for why this method is appropriate in this context. (d) A1 valid confounding variable identified; A1 correct reasoning for why it could confound the results if not controlled/recorded.
Question 4 · Continuous cumulative frequency curve, box plot construction and distribution symmetry
15 marks
The heights, in cm, of 60 plants grown under experimental conditions are grouped below.
(a) Construct a cumulative frequency table for this data. [2] (b) Draw a cumulative frequency diagram, plotting cumulative frequency against the UPPER class boundary of each class. [3] (c) Use your diagram (or interpolation) to estimate the median, the lower quartile and the upper quartile of the plant heights. [4] (d) Hence calculate the interquartile range, and construct a box plot to represent the data (minimum = 0, maximum = 60). [3] (e) Using your quartile values, comment on the skewness of the distribution of plant heights. [3]
Show answer & marking schemeHide answer & marking scheme
(b) A smooth curve (or series of straight-line segments) plotted through the points (10,3), (20,11), (30,26), (40,46), (50,56), (60,60), with height on the horizontal axis and cumulative frequency on the vertical axis.
(c) With n=60: median position = 30th value. This falls in the 30-40 class (cumulative frequency reaches 26 at height 30, and 46 at height 40), so by linear interpolation: \( \text{median} = 30+\dfrac{30-26}{20}\times10 = 30+2.0 = 32.0\ \text{cm} \). LQ position = 15th value, in the 20-30 class (cumulative frequency 11 to 26): \( LQ = 20+\dfrac{15-11}{15}\times10 = 20+2.67 = 22.67\ \text{cm} \). UQ position = 45th value, in the 30-40 class (cumulative frequency 26 to 46): \( UQ = 30+\dfrac{45-26}{20}\times10 = 30+9.5 = 39.5\ \text{cm} \).
(d) \( IQR = UQ-LQ = 39.5-22.67 = 16.83\ \text{cm} \). Box plot: minimum = 0, LQ = 22.67, median = 32.0, UQ = 39.5, maximum = 60, drawn as a box from LQ to UQ (with a line at the median) and whiskers extending to the minimum and maximum values.
(e) The distance from the median to the lower quartile (\( 32.0-22.67=9.33\ \text{cm} \)) is greater than the distance from the median to the upper quartile (\( 39.5-32.0=7.5\ \text{cm} \)); since the lower half of the middle 50% of the data is more spread out than the upper half, the distribution shows a slight negative skew (a longer 'tail' towards the lower height values).
Final answer: median = 32.0 cm, LQ = 22.67 cm, UQ = 39.5 cm, IQR = 16.83 cm; slightly negatively skewed.
Marking scheme
(a) A1 correct upper class boundaries; A1 correct cumulative frequencies (3, 11, 26, 46, 56, 60). (b) A1 correctly labelled axes; A1 all six points plotted accurately; A1 smooth curve or line segments drawn through the points. (c) M1 correct interpolation method shown for at least one quartile; A1 median = 32.0 (accept 31-33); A1 LQ = 22.67 (accept 22-23.5); A1 UQ = 39.5 (accept 39-40). (d) A1 IQR = 16.83 (ECF); A1 box correctly drawn from LQ to UQ with median line; A1 whiskers correctly drawn to given minimum (0) and maximum (60). (e) M1 correct comparison of (median-LQ) with (UQ-median); A1 correct values compared (9.33 vs 7.5); A1 correct conclusion: slight negative skew, with valid reasoning.
Question 5 · Histogram grouped mean estimation and error impact evaluation
6 marks
The weights, in kg, of 50 parcels processed at a depot are grouped below.
(a) Calculate an estimate of the mean weight of a parcel. [3] (b) The depot's records show the TRUE mean weight (calculated from the original, ungrouped data) was actually 11.62 kg. Calculate the error in the grouped-data estimate, and suggest one reason why grouping the data introduces this kind of error. [3]
Show answer & marking schemeHide answer & marking scheme
(b) Error \( = 11.40-11.62 = -0.22\ \text{kg} \) (the grouped-data estimate is 0.22 kg lower than the true mean). This error arises because the calculation assumes every parcel within a class interval weighs exactly the midpoint value, when in reality the individual parcel weights are spread throughout each interval (not concentrated exactly at the midpoint), so some information about the exact distribution of weights within each class is lost when the data is grouped.
Final answer: (a) 11.40 kg; (b) error = -0.22 kg.
Marking scheme
(a) M1 correct method (Σfx/Σf) with correct Σfx = 570 shown; A1 correct mean = 11.40 kg. (b) A1 correct error calculation, -0.22 kg (ECF from (a)); A1 valid reason (midpoint assumption/loss of information about spread within classes) correctly explained.
Question 6 · Comparative pie charts with area-proportional radius computation
8 marks
Two surveys were conducted on preferred transport method: Survey A had 200 respondents, and Survey B had 450 respondents. The results are to be shown as two comparative pie charts, where the AREA of each pie chart is proportional to the number of respondents in that survey.
(a) Explain why, for area-proportional comparative pie charts, the RADIUS of each circle should be proportional to the SQUARE ROOT of the number of respondents (rather than directly proportional to the number of respondents). [2] (b) The pie chart for Survey A is drawn with a radius of 4 cm. Calculate the radius that should be used for the pie chart for Survey B, so that the areas of the two pie charts are correctly proportional to the two sample sizes. [4] (c) In Survey A, 90 out of the 200 respondents preferred public transport. Calculate the angle that should be used to represent this category in the pie chart for Survey A. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) The area of a circle is proportional to the SQUARE of its radius (\( A=\pi r^2 \)); since we want the AREA (not the radius itself) to be proportional to the number of respondents, the radius must be proportional to the SQUARE ROOT of the number of respondents, so that when it is squared (to give the area), the area comes out correctly proportional to the sample size.
(a) A1 correctly identifies area proportional to radius squared (\( A=\pi r^2 \)); A1 correctly explains radius must be proportional to the square root of n for area to be correctly proportional to n. (b) M1 correct method, \( r_B=r_A\sqrt{n_B/n_A} \); A1 correct ratio \( \sqrt{450/200}=1.5 \); A1 \( r_B=6.0\ \text{cm} \). (c) M1 correct method (fraction × 360°); A1 \( 162^\circ \).
Question 7 · Quarterly time series moving average computation, trend plotting and extrapolation assumptions
9 marks
An ice cream parlour's quarterly sales, in £000s, over two years are shown below.
(a) Calculate the uncentred 4-point moving averages for this data. [3] (b) By centring successive pairs of these moving averages, calculate the centred 4-point moving average trend values (which will align with Q3-22, Q4-22, Q1-23 and Q2-23). [3] (c) Comment on the trend shown by your values in (b). [1] (d) The manager wants to use the trend to extrapolate (predict) sales for Q1-24 and beyond. State two assumptions that must be made for this extrapolation to be valid. [2]
Show answer & marking schemeHide answer & marking scheme
(c) The centred moving average trend values (38.0, 39.125, 40.75, 42.125) increase steadily over the four quarters shown, indicating a steady underlying upward trend in sales, once seasonal quarterly variation has been smoothed out.
(d) Two assumptions required: (i) that the underlying (smoothed) trend seen in the data will continue in the same direction and at broadly the same rate into the future; and (ii) that the seasonal pattern of variation between quarters (e.g. higher sales in summer quarters) will remain broadly the same in future years as it was in the data collected.
Marking scheme
(a) M1 correct method shown; A1 first three MAs correct (37.5, 38.5, 39.75); A1 remaining two MAs correct (41.75, 42.5). (b) M1 correct centring method; A1 first two centred values correct (38.0, 39.125); A1 remaining two correct (40.75, 42.125). (c) A1 correctly identifies a steady increasing trend. (d) A1 each for two valid, distinct assumptions (trend continues; seasonal pattern remains the same); max [2].
Question 8 · Three-set Venn diagram, conditional probability and cross-group estimation
13 marks
In a survey of 50 students, F = plays football, B = plays basketball, S = swims. The number of students in each region of a three-set Venn diagram is as follows: football only = 10; basketball only = 8; swimming only = 12; football AND basketball only (not swimming) = 5; football AND swimming only (not basketball) = 4; basketball AND swimming only (not football) = 3; all three activities = 2; none of the three activities = 6.
(a) Draw a three-set Venn diagram (describing all eight regions and their values) to represent this information, and verify that the total is 50. [3] (b) Find the total number of students who play football, and the total number who swim. [2] (c) Find \( P(F\cap B\cap S) \), the probability that a randomly selected student does all three activities. [1] (d) Find \( P(\text{swims} \mid \text{plays football}) \), the conditional probability that a student swims, GIVEN that they play football. [3] (e) Determine whether the events 'plays football' and 'swims' are statistically independent, showing your working. [4]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) Venn diagram regions: football only = 10, basketball only = 8, swimming only = 12, F∩B only = 5, F∩S only = 4, B∩S only = 3, F∩B∩S = 2, none = 6. Sum: \( 10+8+12+5+4+3+2+6=50 \), which matches the total of 50 students surveyed, confirming the diagram is consistent.
(b) \( |F| = 10+5+4+2 = 21 \) students play football. \( |S| = 12+4+3+2 = 21 \) students swim.
(e) For independence, we require \( P(S\cap F) = P(S)\times P(F) \). \( P(S)\times P(F) = \dfrac{21}{50}\times\dfrac{21}{50} = \dfrac{441}{2500} = 0.1764 \). \( P(S\cap F) = \dfrac{6}{50} = 0.12 \). Since \( 0.12 \neq 0.1764 \), the events 'plays football' and 'swims' are NOT statistically independent.
Final answer: (b) |F|=21, |S|=21; (c) 0.04; (d) 2/7 ≈ 0.286; (e) not independent.
Marking scheme
(a) A1 all eight regions correctly placed; A1 correct sum shown (50); A1 diagram correctly structured with three overlapping circles/universal set boundary described. (b) A1 |F|=21; A1 |S|=21. (c) A1 \( P(F\cap B\cap S)=2/50=0.04 \). (d) M1 correct identification of \( P(S\cap F)=6/50 \); M1 correct conditional probability method; A1 \( P(S|F)=2/7\approx0.286 \). (e) M1 correct method, comparing \( P(S\cap F) \) with \( P(S)\times P(F) \); A1 correct value of \( P(S)\times P(F)=0.1764 \); A1 correct value of \( P(S\cap F)=0.12 \); A1 correct conclusion, not independent, since the two values differ.
Question 9 · Chain-base index numbers, missing price computation and geometric mean calculation
10 marks
The price of a raw material was £20.00 in 2020. The chain-base index numbers (each year expressed relative to the PREVIOUS year, ×100) for the following three years were:
Year 2021 2022 2023 Chain-base index 110 105 108
(a) Explain what is meant by a 'chain-base' index number, as opposed to a 'fixed-base' index number. [2] (b) Calculate the price of the raw material in 2021. [2] (c) Calculate the price of the raw material in 2022 (the 'missing' price). [2] (d) Calculate the price of the raw material in 2023. [2] (e) Calculate the geometric mean of the three annual growth factors (1.10, 1.05 and 1.08) implied by the chain-base indices, and interpret this value as an average annual percentage growth rate over the three years. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) A chain-base index compares each period's value with the value in the IMMEDIATELY PRECEDING period (the base changes/updates every year), showing the year-on-year percentage change; a fixed-base index, by contrast, compares every period's value with the value in a single, fixed base year that does not change.
(b) Chain-base index for 2021 = 110 means the 2021 price is 110% of the 2020 price: \( \text{price}_{2021} = 20.00\times\dfrac{110}{100} = £22.00 \)
(c) Chain-base index for 2022 = 105 means the 2022 price is 105% of the 2021 price: \( \text{price}_{2022} = 22.00\times\dfrac{105}{100} = £23.10 \)
(d) Chain-base index for 2023 = 108 means the 2023 price is 108% of the 2022 price: \( \text{price}_{2023} = 23.10\times\dfrac{108}{100} = £24.95 \) (2 d.p.)
(e) Geometric mean \( = \sqrt[3]{1.10\times1.05\times1.08} = \sqrt[3]{1.2474} = 1.0765 \) (4 d.p.). This corresponds to an average annual growth rate of approximately \( (1.0765-1)\times100 = 7.65\% \) — i.e. if the price had grown at this same, constant rate every year instead of the three different actual rates, it would have reached the same final 2023 price from the 2020 starting price.
Final answer: (b) £22.00; (c) £23.10; (d) £24.95; (e) geometric mean growth rate ≈ 7.65% per year.
Marking scheme
(a) A1 correct description of chain-base (compares with previous period); A1 correct description of fixed-base (compares with one fixed base period), clearly distinguishing the two. (b) M1 correct method (price × index/100); A1 £22.00. (c) M1 correct method applied to the 2021 price (ECF); A1 £23.10. (d) M1 correct method applied to the 2022 price (ECF); A1 £24.95. (e) M1 correct method, cube root of the product of the three growth factors; A1 correct geometric mean growth rate ≈7.65% (accept 7.6-7.7%), with correct interpretation as an average annual rate.
Question 10 · Probability tree without replacement and conditional probability deduction
11 marks
A bag contains 5 red marbles and 3 blue marbles. Two marbles are drawn from the bag at random, ONE AFTER THE OTHER, WITHOUT REPLACEMENT (the first marble is not put back before the second is drawn).
(a) Complete a probability tree diagram for the two draws, labelling all branch probabilities as fractions. [4] (b) Calculate the probability that both marbles drawn are red. [2] (c) Calculate the probability that at least one of the two marbles drawn is red. [3] (d) Given that the first marble drawn was blue, state the probability that the second marble drawn is red, and explain why this is different from the probability of drawing a red marble on the first draw. [2]
Show answer & marking schemeHide answer & marking scheme
Worked solution
(a) First draw: P(Red) = 5/8, P(Blue) = 3/8. If the first marble is red (7 marbles remain: 4 red, 3 blue): second draw P(Red|Red) = 4/7, P(Blue|Red) = 3/7. If the first marble is blue (7 marbles remain: 5 red, 2 blue): second draw P(Red|Blue) = 5/7, P(Blue|Blue) = 2/7.
(d) Given the first marble was blue, 7 marbles remain in the bag: 5 red and 2 blue (one blue marble has been removed and not replaced). So \( P(\text{2nd is red}\mid\text{1st is blue}) = \dfrac{5}{7} \). This differs from the probability of drawing red on the first draw (5/8) because, since the marbles are drawn WITHOUT replacement, removing the first (blue) marble changes the total number of marbles remaining in the bag (from 8 to 7) and the composition of the bag, so the probabilities for the second draw are conditional on (dependent on) the outcome of the first draw.
Final answer: (b) 5/14; (c) 25/28; (d) 5/7.
Marking scheme
(a) A1 correct first-draw probabilities (5/8, 3/8); A1 correct second-draw probabilities following red (4/7, 3/7); A1 correct second-draw probabilities following blue (5/7, 2/7); A1 diagram structure correctly shows four branch pairs with denominators changing to 7. (b) M1 correct method; A1 \( P(RR)=5/14 \). (c) M1 correctly identifies \( 1-P(BB) \) as the method; M1 correct calculation of P(BB)=3/28; A1 \( P(\text{at least one red})=25/28 \). (d) A1 correct value, 5/7; A1 correct explanation referencing the changed total/composition of the bag due to sampling without replacement.
Question 11 · Contingency table relative risk analysis and binomial distribution model evaluation
13 marks
A health study recorded the incidence of a respiratory condition among smokers and non-smokers:
(a) Calculate the risk (probability) of the condition for a smoker, and for a non-smoker, expressing each as a percentage. [2] (b) Calculate the relative risk of the condition for smokers compared with non-smokers, and interpret this value in context. [3] (c) Calculate the absolute risk difference (in percentage points) between smokers and non-smokers, and interpret this value in context. [2] (d) Assuming each smoker independently has a 22.5% chance of developing the condition, and that a random sample of 10 smokers is selected, state a suitable probability distribution to model the number of these 10 smokers who develop the condition, including its parameters. [2] (e) Using your model from (d), calculate the probability that EXACTLY 2 of the 10 sampled smokers develop the condition, giving your answer to 3 significant figures. [3] (f) State one assumption of the binomial model used in (d) that may not be entirely realistic in this context, giving a reason. [1]
Show answer & marking schemeHide answer & marking scheme
(b) Relative risk \( = \dfrac{\text{risk (smoker)}}{\text{risk (non-smoker)}} = \dfrac{22.5}{5.0} = 4.5 \). This means a smoker is 4.5 times as likely to develop the condition as a non-smoker.
(c) Absolute risk difference \( = 22.5\%-5.0\% = 17.5 \) percentage points. This means that, for every 100 smokers (compared with 100 non-smokers), approximately 17.5 more people would be expected to develop the condition as a direct/additional result of (association with) smoking.
(d) \( X\sim B(10,\,0.225) \), where X is the number of the 10 sampled smokers who develop the condition, n=10 is the sample size, and p=0.225 is the probability of the condition for each (independent) smoker.
(f) The binomial model assumes each smoker's chance of developing the condition is independent of every other sampled smoker and identical (constant at 22.5%) for each; in reality, this may not be entirely realistic, since individual risk could vary between smokers (e.g. depending on how much/how long they have smoked, or other individual health/genetic factors), meaning the true probability is not identical for every smoker as the model assumes.
(a) A1 risk (smoker) = 22.5%; A1 risk (non-smoker) = 5.0%. (b) M1 correct method (ratio of the two risks); A1 RR = 4.5; A1 correct contextual interpretation (smokers 4.5 times as likely). (c) A1 correct value, 17.5 percentage points; A1 correct contextual interpretation. (d) A1 correctly states \( B(10,0.225) \); A1 correctly identifies/defines both parameters n and p in context. (e) M1 correct binomial coefficient \( \binom{10}{2}=45 \); M1 correct substitution into the binomial formula; A1 \( P(X=2)=0.296 \) (accept 0.294–0.298). (f) A1 valid assumption identified (constant/identical probability, or independence between smokers) with correct reasoning as to why it may not hold exactly in this context.
Wondering how well you actually know this?
thinka is an AI practice app for IGCSE & IB students: unlimited questions, instant auto-marking, and detailed step-by-step solutions. 100,000+ students use it to confirm they actually know it, not just think they do.