CCEA GCSE · thinka 原创模拟试题

2024 CCEA GCSE Statistics 2260 模拟试题及答案详解

Thinka Jun 2024 CCEA GCSE-Style Mock — Statistics 2260

200 240 分钟2024
An original Thinka practice paper modelled on the structure and difficulty of the Jun 2024 CCEA GCSE Statistics 2260 paper. Not affiliated with or reproduced from CCEA.

部分 Unit 1 (With Calculator) Higher Tier

Answer all ten questions. Complete in black ink only. HB pencil may be used for drawing. Show all working clearly. Formula sheet is provided on page 2.
10 题目 · 100
题目 1 · Structured Data Display & Summaries (Stem-Leaf / Box Plots)
9
A researcher records the time, in minutes, taken by 15 students to complete a logic puzzle. The times, already sorted into ascending order, are:
12 15 15 18 21 22 24 24 25 27 29 31 33 36 38

(a) Write the data as an ordered stem-and-leaf diagram, using the tens digit as the stem. Include a key. [3]
(b) Use your diagram to find (i) the median time, (ii) the interquartile range (IQR). [3]
(c) Write down the five-number summary (minimum, lower quartile, median, upper quartile, maximum) for this data, and use it to comment on the skewness of the distribution. [3]
查看答案详解

解题

(a) There are n = 15 values, already ordered. Grouping by tens digit (stem):
1 | 2 5 5 8
2 | 1 2 4 4 5 7 9
3 | 1 3 6 8
Key: 1|2 means 12 minutes.

(b) With n = 15, the median is the 8th value: counting along the ordered list (12,15,15,18,21,22,24,...) the 8th value is 24. Median = 24 minutes.
For the quartiles, using position (n+1)/4 = 16/4 = 4th value for the lower quartile and 3(n+1)/4 = 12th value for the upper quartile: the 4th value is 18 and the 12th value is 31. IQR = 31 − 18 = 13 minutes.

(c) Five-number summary: minimum = 12, lower quartile = 18, median = 24, upper quartile = 31, maximum = 38.
Median − LQ = 24 − 18 = 6; UQ − Median = 31 − 24 = 7. Since UQ − Median (7) is slightly larger than Median − LQ (6), and the gap from median to maximum (14) is larger than from median to minimum (12), the distribution is slightly positively (right) skewed — a few students took noticeably longer than the majority.
Final answer: (a) 1|2 5 5 8, 2|1 2 4 4 5 7 9, 3|1 3 6 8 (key: 1|2 = 12 minutes); (b)(i) median = 24 minutes, (ii) IQR = 13 minutes; (c) five-number summary 12, 18, 24, 31, 38; slightly positively (right) skewed

评分标准

(a) B1 correct stems (1,2,3) with leaves grouped correctly; B1 leaves in each row written in ascending order; B1 correct key with units. [3]
(b)(i) B1 median = 24 (minutes unit not essential here as context given). [1]
(ii) M1 correctly identifies LQ = 18 and UQ = 31 (or equivalent valid quartile method consistently applied); A1 IQR = 13. [2]
(c) B1 correct five-number summary 12, 18, 24, 31, 38; B1 correct numerical comparison of the two quartile gaps (6 vs 7, or min/max gaps); A1 valid conclusion of slight positive/right skew consistent with the figures given. [3]
Accept equivalent valid quartile methods (e.g. Q1 at n/4) provided applied consistently and IQR/skew conclusion follows from the candidate's own figures. Total [9].
题目 2 · Structured Data Display & Summaries (Stem-Leaf / Box Plots)
10
A gym records the number of sessions attended per member, over one month, for two branches. The five-number summaries are:
Branch A: minimum 2, lower quartile 6, median 10, upper quartile 14, maximum 22
Branch B: minimum 4, lower quartile 9, median 11, upper quartile 13, maximum 18

(a) Calculate the interquartile range (IQR) and the range for each branch. [4]
(b) A box plot for each branch would be drawn using these five values. State, with a reason, which branch's box plot would show the more symmetrical distribution. [3]
(c) State, with a reason, which branch has the more consistent (less variable) attendance. [3]
查看答案详解

解题

(a) Branch A: IQR = UQ − LQ = 14 − 6 = 8; range = max − min = 22 − 2 = 20.
Branch B: IQR = 13 − 9 = 4; range = 18 − 4 = 14.

(b) Checking symmetry using the gaps either side of the median:
Branch A: median − LQ = 10 − 6 = 4; UQ − median = 14 − 10 = 4 (equal, so the middle 50% is symmetric), but min-to-median = 10 − 2 = 8 and median-to-max = 22 − 10 = 12 (unequal, so the whiskers are skewed).
Branch B: median − LQ = 11 − 9 = 2; UQ − median = 13 − 11 = 2 (equal); min-to-median = 11 − 4 = 7 and median-to-max = 18 − 11 = 7 (also equal). Branch B has matching gaps on both the box and the whiskers, so its box plot is the more symmetrical of the two.

(c) Branch B has both the smaller IQR (4 vs 8) and the smaller range (14 vs 20), so attendance at Branch B is less spread out and therefore more consistent than at Branch A.
Final answer: (a) Branch A: IQR = 8, range = 20; Branch B: IQR = 4, range = 14; (b) Branch B, since both its box gaps (2 and 2) and its whisker gaps (7 and 7) are equal, whereas Branch A's whisker gaps (8 and 12) are unequal; (c) Branch B is more consistent (smaller IQR and smaller range)

评分标准

(a) M1 IQR method for either branch; A1 Branch A IQR = 8; A1 Branch B IQR = 4; B1 both ranges correct (20 and 14). [4]
(b) M1 compares median−LQ with UQ−median for at least one branch; A1 compares min-to-median with median-to-max for at least one branch; A1 correct conclusion (Branch B) with valid supporting reason from the candidate's figures. [3]
(c) B1 correct conclusion (Branch B); B1 reason referring to smaller IQR; B1 reason referring to smaller range. [3]
Total [10].
题目 3 · Planning & Sampling Theory
6
A school has 900 students: 300 in Year 9, 350 in Year 10 and 250 in Year 11. A researcher wants to take a stratified random sample of 90 students, stratified by year group.

(a) Calculate the number of students that should be sampled from each year group. [3]
(b) Give one advantage of using a stratified sample rather than a simple random sample in this context. [1]
(c) Give one disadvantage of using a stratified sample rather than a simple random sample in this context. [2]
查看答案详解

解题

(a) The sampling fraction is 90/900 = 1/10. Applying this to each year group:
Year 9: 300 × 1/10 = 30; Year 10: 350 × 1/10 = 35; Year 11: 250 × 1/10 = 25.
Check: 30 + 35 + 25 = 90, which matches the required sample size.

(b) Stratified sampling guarantees that each year group is represented in the sample in proportion to its size in the school, so no year group can be over- or under-represented by chance, which a simple random sample cannot guarantee.

(c) Stratified sampling requires a full sampling frame that records which year group (stratum) each of the 900 students belongs to before sampling can begin; compiling and organising the population into strata like this takes extra time and administrative effort compared with simply numbering the whole population for a simple random sample.
Final answer: (a) Year 9: 30, Year 10: 35, Year 11: 25; (b) it ensures proportional representation from every year group, which a simple random sample cannot guarantee; (c) it requires a sampling frame that lists which year group every student belongs to, which takes more time and organisation to set up than a simple random sample

评分标准

(a) M1 correct sampling fraction 90/900 (or equivalent scaling); A1 all three values correct (30, 35, 25); B1 values sum to 90 (can be implied by correct values). [3]
(b) B1 any valid advantage referring to proportional/guaranteed representation of subgroups. [1]
(c) B1 valid disadvantage identified (e.g. needs a sampling frame divided into strata / more time-consuming to set up); B1 further development/explanation of that disadvantage. [2]
Total [6].
题目 4 · Index Numbers & Economic Statistics
11
A single bus ticket in a city cost £2.40 in 2019, £2.52 in 2020, and £2.64 in 2021.

(a) State what is meant by the 'base year' of an index number series. [1]
(b) Using 2019 as the base year (index = 100), calculate the simple price index number for the bus ticket in 2020 and in 2021. [4]
(c) Calculate the chain base index number for 2021, using 2020 as the reference year. [3]
(d) The city's Retail Price Index (RPI), also based at 2019 = 100, stood at 118 in 2021. The bus company claims that, after accounting for inflation, bus fares have 'more than doubled in real terms' since 2019. Use your index number for 2021 and the RPI to evaluate this claim. [3]
查看答案详解

解题

(a) The base year is the year chosen as the point of comparison for an index number series; it is always given an index number of 100, and every other year's index shows its price as a percentage of the base year price.

(b) Simple index = (price in given year / price in base year) × 100.
2020: (2.52 / 2.40) × 100 = 105.
2021: (2.64 / 2.40) × 100 = 110.

(c) Chain base index compares each year to the immediately preceding year:
2021 chain base index = (2.64 / 2.52) × 100 = 104.761... ≈ 104.8 (1 d.p.). This means the ticket price rose by about 4.8% from 2020 to 2021.

(d) To find the real-terms (inflation-adjusted) index, divide the price index by the RPI and multiply by 100:
Real-terms index (2021) = (110 / 118) × 100 = 93.2 (1 d.p.).
Since this real-terms index (93.2) is below 100, bus fares actually fell slightly in real terms between 2019 and 2021, once general inflation is accounted for. The claim that fares 'more than doubled in real terms' is therefore false — a real-terms doubling would require a real-terms index of 200.
Final answer: (a) the year against which all other years are compared, and which is given an index number of 100; (b) 2020: 105, 2021: 110; (c) chain base index for 2021 = 104.8; (d) real-terms index = 110/118 x 100 = 93.2, which is below 100, so fares actually fell in real terms — the claim that fares more than doubled in real terms is false

评分标准

(a) B1 correct definition (reference year, index 100). [1]
(b) M1 correct method (price/base price × 100) applied to 2020; A1 2020 = 105; M1 method applied to 2021; A1 2021 = 110. [4]
(c) M1 correct chain base method (2021 price / 2020 price × 100); A1 104.8 (accept 104.76 to 104.8); B1 correct contextual interpretation (≈4.8% rise 2020 to 2021). [3]
(d) M1 correctly computes real-terms index = 93.2 (accept 93.2 or unrounded 93.22); A1 correctly identifies value is below 100 so a real fall, not a real rise; A1 valid conclusion that the 'more than doubled' claim is false, with reference to what an index of 200 would represent. [3]
Total [11].
题目 5 · Index Numbers & Economic Statistics
12
A trade union monitors the cost of living for its members using three spending categories. The table shows the weight (based on typical spending) and the price index (base year = 100) for each category this year.

Category Weight Price index
Food 5 108
Housing 3 112
Transport 2 104

(a) Explain why a weighted index number may give a more accurate picture of overall price change than a simple (unweighted) index number here. [2]
(b) Calculate the weighted index number for the overall cost of living. [4]
(c) Interpret the weighted index number you found in (b) in the context of the cost of living. [2]
(d) An employee earned £24,000 in the base year. Using your weighted index number, calculate the salary this employee would need to earn this year to maintain the same purchasing power. [3]
(e) State one limitation of using fixed weights, taken from a single base year, in this type of index. [1]
查看答案详解

解题

(a) A simple index would treat Food, Housing and Transport as equally important, but households actually spend different amounts on each. A weighted index multiplies each category's price index by how much is actually spent on it (its weight), so categories that take up a larger share of the household budget have a correspondingly larger effect on the overall figure, giving a more realistic picture of the true change in the cost of living.

(b) Weighted index = Σ(weight × price index) / Σ(weight).
Σ(weight × price index) = (5 × 108) + (3 × 112) + (2 × 104) = 540 + 336 + 208 = 1084.
Σ(weight) = 5 + 3 + 2 = 10.
Weighted index = 1084 / 10 = 108.4.

(c) A weighted index of 108.4 means that, once spending patterns are accounted for, the overall cost of living has risen by 8.4% since the base year.

(d) To maintain the same purchasing power, the salary must rise by the same 8.4%:
New salary = £24,000 × (108.4 / 100) = £24,000 × 1.084 = £26,016.

(e) The weights are fixed at the values recorded in the base year, but household spending patterns change over time (for example, spending on transport or housing as a proportion of income may rise or fall); this means the index becomes progressively less representative of true spending patterns as more time passes since the base year, unless the weights are periodically updated.
Final answer: (a) a weighted index accounts for the fact that households spend more on some categories than others, so it reflects true overall spending patterns rather than treating every category as equally important; (b) weighted index = 108.4; (c) the overall cost of living has risen by 8.4% since the base year, once spending patterns are taken into account; (d) £26,016; (e) fixed weights become outdated as spending habits change over time, so the index becomes less accurate the further it moves from the base year

评分标准

(a) B1 identifies that categories are not equally important / households spend different amounts on each; B1 explains weighting reflects this to give a more accurate overall figure. [2]
(b) M1 correct method Σ(weight × index); A1 Σ(weight × index) = 1084; M1 divides by Σ(weight) = 10; A1 108.4. [4]
(c) B1 correctly reads off 8.4% change; B1 correctly links this to 'cost of living has risen', in context. [2]
(d) M1 correct method (24,000 × 108.4/100, or equivalent); A1 26,016; B1 correct £ units / sensible final answer. [3]
(e) B1 valid limitation referring to weights becoming outdated / spending patterns changing over time. [1]
Total [12].
题目 6 · Probability & Venn Diagrams
8
In a survey of 60 students, 32 play football (F), 25 play rugby (R), and 12 play both football and rugby.

(a) Using a universal set of 60 students, find the number of students in each of the four regions of a two-set Venn diagram for F and R (football only, rugby only, both, neither), and describe where each number would be placed. [3]
(b) Find the probability that a randomly selected student plays neither sport. [1]
(c) Find the probability that a randomly selected student plays football only. [1]
(d) Find the probability that a randomly selected student plays rugby, given that they play football. [3]
查看答案详解

解题

(a) 12 students play both sports, so this goes in the overlap of the two circles.
Football only = total playing football − both = 32 − 12 = 20 (placed in the F circle, outside the overlap).
Rugby only = total playing rugby − both = 25 − 12 = 13 (placed in the R circle, outside the overlap).
Students playing at least one sport = 20 + 13 + 12 = 45, so those playing neither = 60 − 45 = 15 (placed outside both circles, inside the rectangle representing the universal set).
Check: 20 + 13 + 12 + 15 = 60. ✓

(b) P(neither) = 15/60 = 0.25.

(c) P(football only) = 20/60 = 1/3 (≈ 0.333).

(d) P(rugby | football) = P(both) / P(football) = 12/32 = 0.375 (using the conditional probability formula restricted to the 32 students who play football, of whom 12 also play rugby).
Final answer: (a) football only = 20, rugby only = 13, both = 12, neither = 15; (b) 15/60 = 0.25; (c) 20/60 = 1/3; (d) P(R|F) = 12/32 = 0.375

评分标准

(a) M1 both = 12 placed correctly in overlap; A1 football only = 20 and rugby only = 13 both correct; A1 neither = 15 correct and all four values shown to sum to 60. [3]
(b) B1 0.25 (or 15/60, 15%) — follow through from (a). [1]
(c) B1 1/3 (or 20/60, 0.333 to 3sf) — follow through from (a). [1]
(d) M1 correct conditional probability set-up (both/football, i.e. 12/32, not 12/60); A1 correct unsimplified value 12/32; A1 0.375 (or 3/8) — follow through from (a). [3]
Total [8].
题目 7 · Continuous Data Representations (Histograms)
6
The grouped frequency table shows the time, in minutes, spent on homework one evening by a sample of students.

Time (minutes) Frequency Frequency density
0 – 10 8 0.8
10 – 20 22 2.2
20 – 30 30 3.0
30 – 50 24 ?
50 – 80 12 ?

(a) Explain why frequency density, rather than frequency, must be plotted on the vertical axis of a histogram when the class widths are unequal, as they are here. [1]
(b) Calculate the frequency density for the classes 30 – 50 minutes and 50 – 80 minutes. [2]
(c) Using the frequency densities given for the classes 10 – 20 and 20 – 30, estimate the number of students who spent between 15 and 25 minutes on homework. [3]
查看答案详解

解题

(a) In a histogram, it is the area of each bar (not its height) that represents the frequency. If the classes have unequal widths and frequency is plotted directly on the vertical axis, a class with a wide interval would have a tall bar purely because of its width, even if its frequency density (concentration of data) is actually low, giving a misleading picture. Using frequency density = frequency ÷ class width ensures the area of each bar correctly represents its frequency, whatever its width.

(b) Frequency density = frequency ÷ class width.
30 – 50: width = 50 − 30 = 20, so frequency density = 24 ÷ 20 = 1.2.
50 – 80: width = 80 − 50 = 30, so frequency density = 12 ÷ 30 = 0.4.

(c) The interval 15 to 25 minutes cuts across two classes:
15 to 20 minutes is part of the 10–20 class (frequency density 2.2), a width of 5 minutes: 5 × 2.2 = 11 students.
20 to 25 minutes is part of the 20–30 class (frequency density 3.0), a width of 5 minutes: 5 × 3.0 = 15 students.
Total estimate = 11 + 15 = 26 students.
Final answer: (a) using frequency directly would make wider bars look taller purely because of their width, distorting the shape; frequency density (frequency ÷ class width) keeps area proportional to frequency regardless of class width; (b) 30–50: 1.2, 50–80: 0.4; (c) 26 students

评分标准

(a) B1 valid explanation referring to area (not height) representing frequency, or wide classes otherwise appearing disproportionately tall/short. [1]
(b) B1 30–50 frequency density = 1.2; B1 50–80 frequency density = 0.4 (each mark independent, correct method frequency/width implied by correct answer). [2]
(c) M1 correctly splits the 15–25 interval into 15–20 and 20–25 portions; M1 correct partial-area calculation for at least one portion (5 × density); A1 total = 26 students. [3]
Total [6].
题目 8 · Bivariate Correlation (Spearman's Rank)
8
Seven products are ranked by a panel of judges for 'taste' (1 = best taste) and separately ranked by their shelf price (1 = cheapest). The ranks are:

Product A B C D E F G
Taste rank 1 2 3 4 5 6 7
Price rank 3 1 2 5 4 7 6

(a) State why Spearman's rank correlation coefficient, rather than the product moment correlation coefficient (PMCC), is the appropriate statistic to calculate here. [1]
(b) Calculate the value of Spearman's rank correlation coefficient for this data. [5]
(c) Interpret the value you found in (b) in the context of taste and price. [2]
查看答案详解

解题

(a) The data given are already ranks (ordinal positions), not raw numerical measurements on a continuous scale. Spearman's rank correlation coefficient is designed for exactly this situation, whereas the PMCC requires random, continuous ('random on random') measurements.

(b) Using Spearman's formula rs = 1 − 6Σd² / [n(n² − 1)], where d = difference between the two ranks for each product:

Product Taste Price d = Taste − Price d²
A 1 3 −2 4
B 2 1 1 1
C 3 2 1 1
D 4 5 −1 1
E 5 4 1 1
F 6 7 −1 1
G 7 6 1 1

Σd² = 4 + 1 + 1 + 1 + 1 + 1 + 1 = 10.
n = 7, so n(n² − 1) = 7 × (49 − 1) = 7 × 48 = 336.
rs = 1 − (6 × 10)/336 = 1 − 60/336 = 1 − 0.17857... = 0.8214... ≈ 0.821 (3 s.f.).

(c) A value of rs ≈ 0.821 is close to +1, indicating a fairly strong positive correlation between taste rank and price rank: in general, products that were ranked better for taste also tended to be ranked as more expensive (a higher price rank number), though the relationship is not perfect.
Final answer: (a) the data given are ranks, not exact numerical measurements, so Spearman's coefficient (designed for ranked/ordinal data) is appropriate rather than PMCC (designed for random, continuous measurements); (b) rs = 0.821 (3 s.f.); (c) there is a fairly strong positive correlation between taste rank and price rank, suggesting products that rank better for taste also tend to rank as more expensive

评分标准

(a) B1 correctly identifies data are ranks/ordinal, so Spearman's is appropriate (not raw random measurements). [1]
(b) M1 correct differences d for all 7 products; M1 correct d² values; A1 Σd² = 10; M1 correct substitution into rs = 1 − 6Σd²/[n(n²−1)] with n = 7; A1 rs = 0.821 (accept 0.82 or unrounded 0.8214). [5]
(c) B1 correctly identifies a strong/fairly strong positive correlation; B1 valid contextual interpretation linking taste rank and price rank. [2]
Total [8].
题目 9 · Statistical Distributions (Normal & Reliability)
17
A manufacturer tests the lifetime of a batch of light bulbs. The lifetimes are normally distributed with a mean of 500 hours and a standard deviation of 20 hours.

(a) State two properties of the normal distribution that make it a reasonable model for a light bulb's lifetime. [2]
(b) Calculate the standardised score (z-score) for a bulb that lasts 460 hours, and interpret this value. [3]
(c) Using the fact that approximately 95% of the data in a normal distribution lies within two standard deviations of the mean, find the percentage of bulbs with a lifetime between 460 and 540 hours. [2]
(d) In a batch of 2000 bulbs, estimate the number of bulbs with a lifetime of less than 460 hours. [3]
(e) A reliability engineer claims that a lifetime exceeding 560 hours would be 'very unusual' for this batch. Use a standardised score to justify this claim. [3]
(f) A second, cheaper brand of bulb has lifetimes normally distributed with mean 480 hours and standard deviation 15 hours. A bulb of the original brand lasts 530 hours; a bulb of the second brand lasts 505 hours. Using standardised scores, determine which of the two bulbs performed better relative to its own brand. [4]
查看答案详解

解题

(a) The normal distribution is symmetric and bell-shaped about its mean, with values becoming progressively less likely the further they are from the mean; manufactured lifetimes typically cluster around a typical value with random small variations either side, and extreme lifetimes (very short or very long) are rare, matching this shape.

(b) z = (x − μ) / σ = (460 − 500) / 20 = −40 / 20 = −2.0.
This means a lifetime of 460 hours is exactly 2 standard deviations below the mean.

(c) 460 = 500 − 2(20) = μ − 2σ, and 540 = 500 + 2(20) = μ + 2σ. So 460 to 540 hours is exactly the interval within 2 standard deviations of the mean, which contains approximately 95% of the data.

(d) Since 95% of bulbs lie between 460 and 540 hours, the remaining 5% lie outside this range, split equally between the two tails (by symmetry): 5% / 2 = 2.5% lie below 460 hours.
Estimated number of bulbs below 460 hours = 2.5% × 2000 = 0.025 × 2000 = 50 bulbs.

(e) z = (560 − 500) / 20 = 60 / 20 = 3.0.
A lifetime of 560 hours is exactly 3 standard deviations above the mean. Since values more than 3 standard deviations from the mean are described as very unusual for a normal distribution, and 560 hours sits exactly at this threshold, the engineer's claim that exceeding 560 hours would be very unusual is justified.

(f) Standardise each bulb relative to its own brand:
Original brand (μ = 500, σ = 20), lifetime 530: z = (530 − 500)/20 = 30/20 = 1.5.
Second brand (μ = 480, σ = 15), lifetime 505: z = (505 − 480)/15 = 25/15 = 1.667 (3 d.p.).
Since 1.667 > 1.5, the second brand's bulb is further above its own brand's mean, in standard-deviation terms, than the original brand's bulb is above its mean. So the second brand's bulb performed relatively better compared with its own brand's typical performance.
Final answer: (a) it is symmetric/bell-shaped about the mean, and most naturally occurring measurements (like manufactured lifetimes) cluster around a central value with fewer extreme values; (b) z = -2.0, meaning 460 hours is exactly 2 standard deviations below the mean; (c) 95%; (d) 50 bulbs; (e) z = 3.0, which is at the threshold beyond which values are classed as very unusual, so the claim is justified; (f) the second brand's bulb (z = 1.67) performed relatively better than the first brand's bulb (z = 1.5)

评分标准

(a) B1 symmetric/bell-shaped property stated; B1 second valid property (e.g. clustering around mean / extreme values rare), applied to bulb lifetimes. [2]
(b) M1 correct formula z = (x−μ)/σ; A1 z = −2.0; B1 correct interpretation (460 is 2 sd below mean). [3]
(c) M1 recognises 460 and 540 are μ∓2σ; A1 95%. [2]
(d) M1 correctly halves the remaining 5% to get 2.5% in the lower tail; M1 applies 2.5% to 2000; A1 50 bulbs. [3]
(e) M1 correct z = 3.0; A1 recognises this is exactly the 3-sd 'very unusual' threshold; A1 valid conclusion the claim is justified, with reference to the 3-sd rule. [3]
(f) M1 correct z for original brand bulb (1.5); M1 correct z for second brand bulb (1.667, accept 1.67 or 25/15); A1 both z-values correct together; A1 correct conclusion (second brand's bulb performed relatively better) with valid comparison of the two z-values. [4]
Total [17].
题目 10 · Statistical Distributions (Binomial Probability)
13
A supplier states that, on average, 30% of components in a delivery are defective. A quality inspector takes a random sample of 8 components from a delivery. Let X be the number of defective components in the sample.

(a) State two conditions that must hold for X to be modelled by a binomial distribution. [2]
(b) Calculate P(X = 3). [3]
(c) Calculate P(X ≥ 3). [4]
(d) Calculate the mean and the standard deviation of X. [3]
(e) The inspector rejects the whole delivery if 3 or more defective components are found in the sample of 8. Using your answer to (c), comment on how often deliveries would be rejected under this rule if the supplier's 30% defect rate is accurate. [1]
查看答案详解

解题

(a) X ~ B(n, p) requires: a fixed number of trials/components tested (n = 8), each with only two possible outcomes (defective or not defective); a constant probability of a component being defective (p = 0.3) for every component; and the components must be tested independently of one another.

(b) X ~ B(8, 0.3). Using P(X = k) = ⁿCₖ pᵏ (1−p)ⁿ⁻ᵏ:
P(X = 3) = ⁸C₃ (0.3)³ (0.7)⁵ = 56 × 0.027 × 0.16807 = 0.25412... ≈ 0.254 (3 s.f.).

(c) P(X ≥ 3) = 1 − P(X = 0) − P(X = 1) − P(X = 2).
P(X = 0) = ⁸C₀ (0.3)⁰ (0.7)⁸ = 1 × 1 × 0.05765 = 0.05765.
P(X = 1) = ⁸C₁ (0.3)¹ (0.7)⁷ = 8 × 0.3 × 0.08235 = 0.19765.
P(X = 2) = ⁸C₂ (0.3)² (0.7)⁶ = 28 × 0.09 × 0.11765 = 0.29648.
Sum = 0.05765 + 0.19765 + 0.29648 = 0.55177.
P(X ≥ 3) = 1 − 0.55177 = 0.44823 ≈ 0.448 (3 s.f.).

(d) Mean = np = 8 × 0.3 = 2.4.
Variance = np(1−p) = 8 × 0.3 × 0.7 = 1.68.
Standard deviation = √1.68 = 1.2961... ≈ 1.30 (3 s.f.).

(e) Since P(X ≥ 3) ≈ 0.448, a delivery would be rejected on approximately 45% of inspections — nearly one in every two deliveries — if the supplier's stated 30% defect rate is accurate. This suggests the rejection rule is triggered very frequently, and the inspector may need either a stricter supplier defect standard or a different sampling/rejection rule.
Final answer: (a) a fixed number of trials/components each with only two outcomes (defective/not defective), and a constant probability of a component being defective, with components tested independently; (b) P(X=3) = 0.254 (3 s.f.); (c) P(X≥3) = 0.448 (3 s.f.); (d) mean = 2.4, standard deviation ≈ 1.30 (3 s.f.); (e) deliveries would be rejected on nearly half of all inspections (about 45%), so the rule triggers rejection very frequently if the 30% defect rate is accurate

评分标准

(a) B1 fixed number of trials with two outcomes; B1 constant probability and independence. [2]
(b) M1 correct binomial formula/set-up with n=8, p=0.3, k=3; M1 correct combinatorial coefficient ⁸C₃ = 56; A1 P(X=3) = 0.254 (accept 0.2541). [3]
(c) M1 correct complement method 1 − P(0) − P(1) − P(2); M1 P(X=0) and P(X=1) both correct (0.0576, 0.1977 to 3sf, or equivalent unrounded); M1 P(X=2) correct (0.2965 to 3sf); A1 P(X≥3) = 0.448. [4]
(d) M1 mean = np = 2.4; M1 variance = np(1−p) = 1.68; A1 sd = 1.30 (accept 1.296). [3]
(e) B1 valid comment quantifying the frequency (≈45%, nearly half) and a sensible implication for the rejection rule. [1]
Total [13].

准备好测试自己了吗?

将这些笔记转化为考试练习。获取此课题的无限AI题目,即时批改及详细解析。

练习此课题

部分 Unit 2 (With Calculator) Higher Tier

Answer all ten questions. Complete in black ink only. HB pencil may be used for drawing. Show all working clearly. Formula sheet is provided on page 2.
10 题目 · 100
题目 1 · Chart Interpretation & Graphical Distortion Critiques
4
A local company's annual report includes a bar chart of its yearly profit (£ millions) for 2020-2023. The vertical axis is described as starting at £80 million (not £0 million) and rising to £90 million, with the bars for 82, 84, 87 and 88 shown for the four years.

(a) Identify one way in which this chart is misleading. [2]
(b) Explain the effect this has on how a reader is likely to interpret the company's profit growth. [2]
查看答案详解

解题

(a) The vertical (profit) axis does not start at zero — it is truncated to start at £80 million. This is a form of graphical misrepresentation because it removes the visual baseline that would normally let readers judge bar heights fairly.

(b) With the axis starting at £80 million instead of £0, the difference between the shortest bar (82) and the tallest bar (88) is exaggerated: visually the tallest bar looks several times the height of the shortest, when in reality the increase is only from £82m to £88m — a rise of (88−82)/82 × 100 ≈ 7.3%, which is a modest change. A reader glancing at the chart is likely to conclude the company's profit grew dramatically, when the true percentage increase is fairly small.
Final answer: (a) the vertical axis is truncated, starting at £80 million instead of £0; (b) this makes a genuinely small change in profit (£82m to £88m, a rise of about 7%) look like a huge, dramatic increase, because the bars appear several times taller than they should relative to each other

评分标准

(a) B1 identifies the truncated/non-zero vertical axis; B1 states the starting value (£80m) or that it does not start at zero, explicitly. [2]
(b) B1 states that differences/growth are visually exaggerated; B1 develops this with a valid contextual explanation (e.g. actual % change is small, or bars look disproportionately different in height). [2]
Total [4].
题目 2 · Chart Interpretation & Graphical Distortion Critiques
4
A marketing report uses a 3-dimensional pie chart to show the market share of four competing supermarket chains.

(a) Identify one issue with using a 3D pie chart to represent this data. [2]
(b) Suggest an appropriate alternative way of representing this data, and explain why it would give a clearer picture. [2]
查看答案详解

解题

(a) Adding a 3D perspective effect to a pie chart distorts the apparent size of each sector: because of the tilted viewing angle, slices positioned towards the front of the chart appear larger than slices of the same true percentage positioned towards the back or sides, making the true relative market shares hard to compare accurately.

(b) A standard flat (2D) pie chart, or alternatively a bar chart, would give a clearer picture. In a 2D pie chart, each sector's angle is drawn exactly proportional to its share with no perspective distortion, so sectors of equal size look genuinely equal; a bar chart would let the reader compare the four market shares directly using bar height, which the eye judges more accurately than areas or angles viewed in perspective.
Final answer: (a) the 3D perspective distorts the apparent size of the sectors, so slices nearer the front of the chart look larger than slices of the same true value at the back or sides; (b) a standard 2D pie chart (or a bar chart) would be clearer, because sector angles/areas (or bar heights) can then be compared accurately without being distorted by a false sense of depth or perspective

评分标准

(a) B1 identifies perspective/3D effect distorts apparent sector size; B1 explains front vs back/side sectors appear unequal even if true values are equal. [2]
(b) B1 valid alternative chart named (2D pie chart or bar chart); B1 valid explanation of why it is clearer (no perspective distortion / easier to compare heights or true angles). [2]
Total [4].
题目 3 · Data Analysis & Outlier Rules
11
The weekly sales figures, in £'00s, for nine small shops in a retail park are:
12 14 15 15 16 17 18 19 45

(a) Find the median, the lower quartile (Q1) and the upper quartile (Q3) of this data. [3]
(b) Calculate the interquartile range (IQR). [1]
(c) An outlier is defined as any value more than 1.5 × IQR below Q1 or above Q3. Find the lower and upper boundaries for outliers, and use them to identify any outlier(s) in this data. [4]
(d) Calculate the mean of the data (i) including the outlier, (ii) excluding the outlier. [2]
(e) Comment on the effect the outlier has on the mean, and state which measure of central tendency (mean or median) better describes this data set. [1]
查看答案详解

解题

(a) With n = 9, ordered data: 12, 14, 15, 15, 16, 17, 18, 19, 45. The median is the middle (5th) value = 16.
The lower half (below the median, first 4 values) is 12, 14, 15, 15; Q1 is the median of these = (14+15)/2 = 14.5.
The upper half (above the median, last 4 values) is 17, 18, 19, 45; Q3 is the median of these = (18+19)/2 = 18.5.

(b) IQR = Q3 − Q1 = 18.5 − 14.5 = 4.

(c) Lower boundary = Q1 − 1.5 × IQR = 14.5 − 1.5(4) = 14.5 − 6 = 8.5.
Upper boundary = Q3 + 1.5 × IQR = 18.5 + 1.5(4) = 18.5 + 6 = 24.5.
No values in the data are below 8.5, so there are no low outliers. The value 45 is above 24.5, so 45 is an outlier.

(d)(i) Mean including outlier = (12+14+15+15+16+17+18+19+45)/9 = 171/9 = 19.0.
(ii) Excluding the outlier (45), the remaining 8 values sum to 171 − 45 = 126. Mean excluding outlier = 126/8 = 15.75.

(e) Including the single outlier of 45 raises the mean from 15.75 to 19.0 — an increase of 3.25, even though only one of the nine shops actually has sales that high. This shows the mean is heavily influenced by extreme values. The median (16) is unaffected by the size of the outlier (it would stay 16 even if the highest value were far larger or smaller, as long as it remained the largest), so the median is the more representative measure of a 'typical' shop's sales for this data set.
Final answer: (a) median = 16, Q1 = 14.5, Q3 = 18.5; (b) IQR = 4; (c) boundaries 8.5 and 24.5, so 45 is an outlier (no low outliers); (d)(i) mean including outlier = 19.0, (ii) mean excluding outlier = 15.75; (e) the outlier pulls the mean up substantially (from 15.75 to 19.0), so the median (16) is a better, more representative measure for this data set since it is not affected by the extreme value

评分标准

(a) B1 median = 16; B1 Q1 = 14.5; B1 Q3 = 18.5. [3]
(b) B1 IQR = 4 (follow through from (a)). [1]
(c) M1 lower boundary = 8.5; M1 upper boundary = 24.5; A1 correctly states no values below lower boundary; A1 correctly identifies 45 as the (only) outlier above the upper boundary. [4]
(d) B1 mean including outlier = 19.0; B1 mean excluding outlier = 15.75. [2]
(e) B1 valid comment that the outlier inflates the mean and the median is more representative/robust to outliers, with reference to the candidate's own figures. [1]
Total [11].
题目 4 · Time Series & Moving Average Forecasting
7
An ice-cream kiosk records its quarterly sales, in £'00s, over two years:

Year 1: Q1 = 42, Q2 = 68, Q3 = 84, Q4 = 30
Year 2: Q1 = 48, Q2 = 76, Q3 = 92, Q4 = 36

(a) Explain why a 4-point moving average (rather than a 3-point moving average) is the appropriate choice for smoothing this quarterly data. [1]
(b) Calculate the centred 4-point moving averages corresponding to Q4 of Year 1 and Q1 of Year 2. [3]
(c) Given that the centred moving average corresponding to Q3 of Year 1 is 56.75, calculate the seasonal variation for Q3 of Year 1 (actual value minus trend value), and interpret this value in context. [3]
查看答案详解

解题

(a) The data has a clear yearly cycle of 4 quarters (seasons). A 4-point moving average averages exactly one complete cycle every time it is calculated, so the seasonal (quarterly) fluctuations cancel out and only the underlying trend remains. A 3-point average would not span a whole cycle and so would not fully remove the seasonal pattern.

(b) First calculate the (uncentred) 4-point moving averages, each covering 4 consecutive quarters:
A = (Q1Y1+Q2Y1+Q3Y1+Q4Y1)/4 = (42+68+84+30)/4 = 224/4 = 56.0 [sits between Q2Y1 and Q3Y1]
B = (Q2Y1+Q3Y1+Q4Y1+Q1Y2)/4 = (68+84+30+48)/4 = 230/4 = 57.5 [sits between Q3Y1 and Q4Y1]
C = (Q3Y1+Q4Y1+Q1Y2+Q2Y2)/4 = (84+30+48+76)/4 = 238/4 = 59.5 [sits between Q4Y1 and Q1Y2]
D = (Q4Y1+Q1Y2+Q2Y2+Q3Y2)/4 = (30+48+76+92)/4 = 246/4 = 61.5 [sits between Q1Y2 and Q2Y2]
Each raw average sits between two quarters, so consecutive pairs are averaged to centre them on a single quarter:
Centred MA for Q4 Yr1 = (B + C)/2 = (57.5 + 59.5)/2 = 117.0/2 = 58.5
Centred MA for Q1 Yr2 = (C + D)/2 = (59.5 + 61.5)/2 = 121.0/2 = 60.5

(c) The trend (centred moving average) for Q3 Year 1 is given as 56.75. The actual sales figure for Q3 Year 1 is 84.
Seasonal variation = actual − trend = 84 − 56.75 = 27.25.
This positive seasonal variation means that, on average, sales in Q3 (the summer quarter) are about 27.25 (i.e. £2725) higher than the underlying trend would predict, reflecting the seasonal boost in demand for ice cream during the summer months.
Final answer: (a) there are 4 quarters (seasons) in a year, so a 4-point moving average averages exactly one full cycle each time, removing the seasonal effect; (b) Q4 Year 1 trend = 58.5, Q1 Year 2 trend = 60.5; (c) seasonal variation = +27.25, meaning Q3 sales are typically about 27.25 (£2725) above the underlying trend, reflecting higher demand for ice cream in the summer quarter

评分标准

(a) B1 valid explanation referring to 4 quarters/seasons per cycle matching the 4-point average. [1]
(b) M1 correctly calculates at least 3 of the 4 uncentred 4-point moving averages (A, B, C, D); M1 correct centring method (averaging consecutive pairs); A1 both centred values correct (58.5 and 60.5). [3]
(c) M1 correct method: actual − trend; A1 27.25; B1 valid contextual interpretation referencing the size and positive sign of the seasonal variation (summer boost). [3]
Total [7].
题目 5 · Time Series & Moving Average Forecasting
8
Using the same ice-cream kiosk data as the previous question, the centred moving averages (trend values) are: Q3 Year 1 = 56.75, Q4 Year 1 = 58.5, Q1 Year 2 = 60.5, Q2 Year 2 = 62.25.

(a) Calculate the seasonal variation for Q1 of Year 2, given that actual sales in Q1 Year 2 were 48. [2]
(b) Assuming the trend continues to increase at the same constant rate as it did from Q4 Year 1 (58.5) to Q1 Year 2 (60.5), estimate the trend value for Q3 of Year 2. [3]
(c) Using your trend estimate from (b) and the seasonal variation for Q3 (+27.25, from the previous question), forecast the sales for Q3 of Year 2. [2]
(d) State one limitation of using this method to forecast sales for Q3 of Year 2. [1]
查看答案详解

解题

(a) Seasonal variation = actual − trend = 48 − 60.5 = −12.5.

(b) The trend rises from 58.5 (Q4 Year 1) to 60.5 (Q1 Year 2), a step of one quarter. Rate of increase per quarter = 60.5 − 58.5 = 2.0.
From Q1 Year 2 to Q3 Year 2 is a further 2 quarters, so the trend is projected to rise by a further 2 × 2.0 = 4.0.
Estimated trend for Q3 Year 2 = 60.5 + 4.0 = 64.5.

(c) Forecast = trend + seasonal variation = 64.5 + 27.25 = 91.75.
So the forecast for Q3 Year 2 sales is 91.75 (in £'00s), i.e. £9,175.

(d) This forecast relies on the assumption that the underlying trend continues to increase at exactly the same constant rate seen between Q4 Year 1 and Q1 Year 2, and that the seasonal variation for Q3 stays exactly the same (+27.25) as previously calculated. In reality, factors such as weather, new competitors, or changing costs could cause the trend or the seasonal pattern to change, making the forecast unreliable — it becomes less reliable the further ahead it is projected.
Final answer: (a) seasonal variation = -12.5; (b) trend estimate for Q3 Year 2 = 64.5; (c) forecast sales for Q3 Year 2 = 91.75 (£9,175); (d) the forecast assumes the trend keeps rising at exactly the same constant rate and that the seasonal pattern stays exactly the same, which may not hold true in practice (e.g. weather, competition or costs could change)

评分标准

(a) M1 correct method actual − trend; A1 −12.5. [2]
(b) M1 correctly finds the rate of increase per quarter (2.0) from the given trend values; M1 correctly projects this rate forward 2 quarters; A1 64.5. [3]
(c) M1 correct method trend + seasonal variation; A1 91.75 (accept 91.75 or £9,175, follow through from (b)). [2]
(d) B1 valid limitation referring to the assumption of a constant trend rate and/or unchanging seasonal pattern, or forecasting further ahead being less reliable. [1]
Total [8].
题目 6 · Bivariate Data, PMCC & Anomaly Impact
15
A seaside kiosk records the number of hours of sunshine (x) and its ice-lolly sales, in £'0s, (y) on 6 days:

Day 1 2 3 4 5 6
Sunshine (x) 2 4 6 8 10 12
Sales (y) 12 20 28 36 44 20

(a) Using Sxy = Σxy − (ΣxΣy)/n, Sxx = Σx² − (Σx)²/n and Syy = Σy² − (Σy)²/n, calculate Sxy, Sxx and Syy for all 6 days. [4]
(b) Hence calculate the product moment correlation coefficient (PMCC) for all 6 days. [3]
(c) It is later discovered that on Day 6 there was a power cut at the kiosk, meaning sales were unusually low despite the high number of sunshine hours. Recalculate Sxy, Sxx, Syy and the PMCC excluding Day 6. [4]
(d) Comment on the effect that removing Day 6 has on the PMCC, and explain what this suggests about the true underlying relationship between sunshine hours and sales. [2]
(e) Using the pattern shown by the remaining 5 days, predict the sales if there were 14 hours of sunshine, and state whether this prediction would be an example of interpolation or extrapolation. [2]
查看答案详解

解题

(a) n = 6. Σx = 2+4+6+8+10+12 = 42. Σy = 12+20+28+36+44+20 = 160.
Σxy = (2×12)+(4×20)+(6×28)+(8×36)+(10×44)+(12×20) = 24+80+168+288+440+240 = 1240.
Σx² = 4+16+36+64+100+144 = 364. Σy² = 144+400+784+1296+1936+400 = 4960.
Sxy = 1240 − (42×160)/6 = 1240 − 6720/6 = 1240 − 1120 = 120.
Sxx = 364 − 42²/6 = 364 − 1764/6 = 364 − 294 = 70.
Syy = 4960 − 160²/6 = 4960 − 25600/6 = 4960 − 4266.67 = 693.33 (2 d.p.).

(b) PMCC r = Sxy / √(Sxx × Syy) = 120 / √(70 × 693.33) = 120 / √48533.33 = 120 / 220.30 = 0.5447 ≈ 0.545 (3 s.f.).

(c) Excluding Day 6: x = 2,4,6,8,10 and y = 12,20,28,36,44. n = 5.
Σx = 30, Σy = 140, Σxy = (2×12)+(4×20)+(6×28)+(8×36)+(10×44) = 24+80+168+288+440 = 1000.
Σx² = 4+16+36+64+100 = 220. Σy² = 144+400+784+1296+1936 = 4560.
Sxy = 1000 − (30×140)/5 = 1000 − 4200/5 = 1000 − 840 = 160.
Sxx = 220 − 30²/5 = 220 − 900/5 = 220 − 180 = 40.
Syy = 4560 − 140²/5 = 4560 − 19600/5 = 4560 − 3920 = 640.
PMCC = 160 / √(40 × 640) = 160 / √25600 = 160 / 160 = 1.00.

(d) Removing the anomalous Day 6 raises the PMCC from 0.545 (a moderate positive correlation) to exactly 1.00 (a perfect positive correlation). This shows that Day 6 — with unusually low sales despite high sunshine, due to the power cut — was distorting the overall correlation. On a normal day, sunshine hours and ice-lolly sales are, in fact, perfectly linearly related; the power cut on Day 6 was an external factor unrelated to sunshine that broke this pattern for that one day.

(e) Using the 5 non-anomalous days, the relationship is a perfect straight line. The gradient is Sxy/Sxx = 160/40 = 4, and the line passes through the mean point (mean x = 30/5 = 6, mean y = 140/5 = 28), so the equation is y − 28 = 4(x − 6), i.e. y = 4x + 4.
At x = 14: y = 4(14) + 4 = 56 + 4 = 60.
Since 14 hours of sunshine is outside the observed range of x-values used to establish the pattern (2 to 12 hours), this prediction is an example of extrapolation, not interpolation, and should be treated with caution.
Final answer: (a) Sxy = 120, Sxx = 70, Syy = 693.33 (2 d.p.); (b) PMCC = 0.545 (3 s.f.); (c) excluding Day 6: Sxy = 160, Sxx = 40, Syy = 640, PMCC = 1.00; (d) removing Day 6 raises the PMCC from a moderate 0.545 to a perfect 1.00, showing Day 6 was an anomaly that weakened an otherwise perfect linear relationship, and that sunshine hours and sales are, in fact, perfectly linearly related on normal days; (e) predicted sales = 60 (£600); this is extrapolation, since 14 hours is outside the range of the observed x-values (2 to 12)

评分标准

(a) M1 correct Σx, Σy, Σxy; M1 correct Σx², Σy²; A1 Sxy = 120 and Sxx = 70 both correct; A1 Syy = 693.33 (or unrounded 693.3 recurring). [4]
(b) M1 correct PMCC formula substitution; A1 unrounded/intermediate value (e.g. √48533.3 ≈ 220.3); A1 r = 0.545. [3]
(c) M1 correctly excludes Day 6 and recomputes Σx, Σy, Σxy, Σx², Σy² for 5 points; A1 Sxy = 160, Sxx = 40, Syy = 640 all correct; M1 correct PMCC formula substitution; A1 r = 1.00. [4]
(d) B1 correctly compares the two PMCC values and identifies Day 6 as the cause of the weaker correlation; B1 valid conclusion that the true underlying relationship (excluding the anomaly) is perfectly linear. [2]
(e) M1 correct method using gradient/line from the 5-point data (or equivalent valid method) to predict y at x=14; A1 60, AND correctly identifies extrapolation with valid reason (x=14 outside data range 2–12). [2]
Total [15].
题目 7 · Index Number Calculations & Trend Inference
13
The average house price in a town was £140,000 in 2015, £161,000 in 2018, and £175,000 in 2021.

(a) Using 2015 as the base year (index = 100), calculate the simple price index number for average house prices in 2018 and in 2021. [4]
(b) Calculate the percentage increase in the index number from 2018 to 2021, and hence state the percentage increase in average house prices over this period. [3]
(c) A local newspaper claims that 'house prices more than doubled between 2015 and 2021.' Using your index number for 2021, evaluate this claim. [2]
(d) The National House Price Index, also based at 2015 = 100, stood at 121 in 2021. Compare the town's index (from part (a)) to the national index, and interpret what this suggests about the town's housing market relative to the national trend. [3]
(e) Explain one reason why using index numbers, rather than raw prices, is useful when comparing price trends across different regions. [1]
查看答案详解

解题

(a) Simple index = (price in given year / price in base year) × 100.
2018: (161,000 / 140,000) × 100 = 115.
2021: (175,000 / 140,000) × 100 = 125.

(b) Percentage increase in the index from 2018 to 2021 = (125 − 115)/115 × 100 = 10/115 × 100 = 8.6957 ≈ 8.7% (1 d.p.).
Since the index is directly proportional to price, this is also the percentage increase in average house prices over the same period: prices rose by about 8.7% from 2018 to 2021 (check directly: (175,000−161,000)/161,000 × 100 = 8.70%, which agrees).

(c) An index number of 125 in 2021 (base 2015 = 100) means house prices rose by 25% between 2015 and 2021, not by 100% or more. For prices to have 'more than doubled,' the 2021 index would need to be at least 200. Since 125 is far below 200, the newspaper's claim is false.

(d) The town's price index in 2021 was 125, compared with a national index of 121. Since both indices use the same 2015 base of 100, they are directly comparable: the town's prices rose by 25% while the national average rose by 21%. This suggests that house prices in this town grew somewhat faster than the national average over this period.

(e) Index numbers rescale prices from any starting level onto the same base-100 scale, so that regions with very different raw price levels (for example, a town with an average price of £140,000 versus a city with a much higher average price) can be compared directly in terms of their percentage change over time, which raw prices alone cannot show.
Final answer: (a) 2018: 115, 2021: 125; (b) percentage increase 2018 to 2021 = 8.7% (1 d.p.), so average house prices rose by about 8.7% over this period; (c) false — an index of 125 means prices rose by 25% since 2015, not more than doubled (which would need an index of 200 or more); (d) the town's index (125) is higher than the national index (121), so house prices in this town rose faster (25%) than the national average (21%) between 2015 and 2021; (e) index numbers convert prices in different regions (which may start from very different price levels) onto the same base-100 scale, making relative rates of change directly comparable

评分标准

(a) M1 correct method applied to 2018; A1 2018 = 115; M1 correct method applied to 2021; A1 2021 = 125. [4]
(b) M1 correct percentage-change method on the index values; A1 8.7% (accept 8.70% or unrounded 8.6957); B1 correctly states this equals the percentage rise in actual house prices over the same period. [3]
(c) B1 correctly identifies claim as false; B1 valid justification referencing what an index of 200 would represent, or the true 25% rise. [2]
(d) M1 correctly compares 125 to 121 (both meaningful as % rises, 25% vs 21%); A1 correct numerical comparison stated; A1 valid conclusion that town prices rose faster than the national average. [3]
(e) B1 valid reason referring to a common/rescaled base allowing regions with different starting price levels to be compared. [1]
Total [13].
题目 8 · Comparative Pie Charts & Risk Calculations
12
A school's total budget, split between Staffing, Resources and Maintenance, was £800,000 in 2020 and £1,000,000 in 2023. Two comparative pie charts are drawn, one for each year, where the radius of each pie chart is proportional to the square root of the total budget it represents. The 2020 pie chart has a radius of 4 cm.

(a) Calculate the radius that should be used for the 2023 pie chart, to 3 significant figures. [3]
(b) Explain why the radius (rather than the area) is scaled using the square root of the ratio of the budgets. [1]
(c) In the 2020 pie chart, the sector for Staffing has an angle of 162°. Calculate the amount of the 2020 budget spent on Staffing. [3]

A clinical trial compares two groups of 400 patients each. In Group A (given a new drug), 20 patients experienced a side effect. In Group B (given a placebo), 50 patients experienced the same side effect.

(d) Calculate the absolute risk of the side effect in Group A and in Group B. [2]
(e) Calculate the relative risk of the side effect for Group A compared with Group B, and interpret this value in context. [3]
查看答案详解

解题

(a) Since area ∝ budget and area ∝ radius², the radius must scale with the square root of the budget ratio.
Ratio of budgets = 1,000,000 / 800,000 = 1.25.
Radius ratio = √1.25 = 1.1180 (4 d.p.).
2023 radius = 4 × 1.1180 = 4.4721... ≈ 4.47 cm (3 s.f.).

(b) The total budget is represented by the AREA of the pie chart, not by its radius directly. Since the area of a circle is πr², area is proportional to the square of the radius. So, for the chart's area to scale correctly in proportion to the budget, the radius must be scaled by the square root of the budget ratio (if the radius itself were scaled directly with the budget ratio, the area — and so the visual impression of size — would be scaled by the square of the budget ratio, which would be misleading).

(c) The Staffing sector takes up 162° out of the full 360° circle, so it represents a fraction 162/360 of the total 2020 budget.
Amount spent on Staffing = (162/360) × £800,000 = 0.45 × £800,000 = £360,000.

(d) Absolute risk in Group A = 20/400 = 0.05 = 5%.
Absolute risk in Group B = 50/400 = 0.125 = 12.5%.

(e) Relative risk (Group A relative to Group B) = risk in Group A ÷ risk in Group B = 0.05 / 0.125 = 0.4.
A relative risk of 0.4 (less than 1) means that patients in Group A (on the new drug) were only 0.4 times as likely to experience the side effect as those in Group B (on the placebo) — equivalently, the new drug was associated with a 60% lower risk of the side effect compared with the placebo.
Final answer: (a) 4.47 cm (3 s.f.); (b) the AREA of a pie chart (not its radius) represents the total budget, and area is proportional to radius squared, so the radius must scale with the square root of the budget ratio for the area to scale with the budget itself; (c) £360,000; (d) Group A risk = 5%, Group B risk = 12.5%; (e) relative risk = 0.4, meaning patients given the new drug were 0.4 times as likely (60% less likely) to experience the side effect compared with those given the placebo

评分标准

(a) M1 correct budget ratio 1,000,000/800,000 = 1.25; M1 correct square root method (√1.25) applied to the radius; A1 4.47 cm (3 s.f., accept 4.472). [3]
(b) B1 valid explanation referring to area (not radius) representing the budget, and area ∝ radius². [1]
(c) M1 correct fraction 162/360; M1 correctly multiplies by £800,000; A1 £360,000 with correct units. [3]
(d) B1 Group A = 5% (or 0.05); B1 Group B = 12.5% (or 0.125). [2]
(e) M1 correct relative risk method (risk A ÷ risk B); A1 0.4; A1 valid contextual interpretation (less likely / 60% lower risk on the new drug). [3]
Total [12].
题目 9 · Multivariate Analysis & Linear Regression Choice
10
A researcher models the relationship between hours of revision (x) and exam score (y, as a percentage) using the regression equation y = 42 + 3.5x, based on data for students who revised between 2 and 12 hours.

(a) Using the regression equation, estimate the exam score for a student who revises for 8 hours. [2]
(b) Explain why this estimate in (a) is an example of interpolation. [1]
(c) Use the equation to estimate the score for a student who revises for 20 hours, and explain why this particular estimate should be treated with caution. [2]
(d) The researcher notices that, for very high revision times, the scatter of points curves rather than continuing in a straight line (suggesting diminishing returns from extra revision). Explain why the product moment correlation coefficient (PMCC) might not be the most appropriate measure of association for the full data set in this case, and suggest an alternative approach. [3]
(e) The data set includes one student who revised for very few hours but has an unusually high natural aptitude, scoring far above the value the line would predict for their revision time. State whether this student's result would tend to increase or decrease the strength of the PMCC for the data set as a whole, giving a reason. [2]
查看答案详解

解题

(a) Substituting x = 8 into y = 42 + 3.5x: y = 42 + 3.5(8) = 42 + 28 = 70. Estimated score = 70%.

(b) The value x = 8 hours lies within the range of x-values (2 to 12 hours) that was actually used to construct the regression equation, so estimating y within this range is interpolation.

(c) Substituting x = 20: y = 42 + 3.5(20) = 42 + 70 = 112. An estimated score of 112% is impossible, since exam scores cannot exceed 100%. This happens because x = 20 hours lies far outside the range of data (2 to 12 hours) used to build the model — this is extrapolation, and there is no guarantee the same linear pattern continues to hold outside the range where it was observed (indeed, the impossible result shows that it clearly does not).

(d) The PMCC only measures the strength of a LINEAR association between two variables. If the true underlying relationship is curved (as suggested here, with diminishing returns at high revision times), then fitting a straight line and calculating a PMCC would not properly describe the pattern in the data — a high PMCC could still be reported for the linear part of the data while missing the curvature, or a modest PMCC could understate a strong but non-linear relationship. A scatter diagram should first be examined visually; if a curved pattern is confirmed, a non-linear (curved) model, or Spearman's rank correlation coefficient applied to the ranks (which is not restricted to linear relationships), would be a more appropriate approach.

(e) This student's point lies a long way from the line of best fit (well above it, for a low revision time), so it increases the overall scatter of the points around the line. This would decrease (weaken) the strength of the PMCC for the data set as a whole, because the PMCC measures how closely points cluster around a straight line, and this point does not follow the general pattern shown by the rest of the data.
Final answer: (a) 70%; (b) 8 hours is within the range of x-values (2 to 12) used to build the model, so this is interpolation; (c) 112%, which is impossible since exam scores cannot exceed 100% — because 20 hours is outside the data range (2 to 12), this is extrapolation and the linear pattern may not continue to hold; (d) PMCC measures the strength of a LINEAR relationship, but the true relationship curves for high revision times, so a high or low PMCC would not properly capture a curved pattern — a scatter diagram should be examined, and/or a non-linear (curved) model or Spearman's rank correlation could be considered instead; (e) it would decrease the strength (weaken) the PMCC, because this point does not fit the general linear pattern followed by the rest of the data, increasing the scatter around the line of best fit

评分标准

(a) M1 correct substitution into the equation; A1 70%. [2]
(b) B1 correctly explains x=8 lies within the original data range (2–12). [1]
(c) B1 112% correctly calculated; B1 correctly identifies this as extrapolation with valid reason (outside data range, result impossible/unreliable). [2]
(d) B1 correctly states PMCC only measures linear association; B1 correctly links this to the curved pattern described being poorly captured by a straight-line/PMCC approach; B1 valid alternative suggested (examine scatter diagram / fit non-linear model / use Spearman's rank). [3]
(e) B1 correctly states the PMCC would decrease/weaken; B1 valid reason referring to the point increasing scatter around the line / not following the general pattern. [2]
Total [10].
题目 10 · Grouped Continuous Data, Skewness & Linear Interpolation
16
The grouped frequency table shows the times, in minutes, taken by 50 runners to complete a fun run.

Time (minutes) Frequency
0 – 10 4
10 – 20 10
20 – 30 18
30 – 40 12
40 – 50 6

(a) Write down the modal class. [1]
(b) Calculate an estimate of the mean time, using the midpoint of each class. [4]
(c) Using linear interpolation, estimate the median time. [4]
(d) Using linear interpolation, estimate the lower quartile (Q1) and the upper quartile (Q3), and hence calculate an estimate of the interquartile range. [4]
(e) Use your median, Q1 and Q3 to calculate the quartile coefficient of skewness, ( Q3 + Q1 − 2×Median ) / ( Q3 − Q1 ), and interpret its sign in the context of the run times. [3]
查看答案详解

解题

(a) The class with the highest frequency (18) is 20 – 30 minutes, so this is the modal class.

(b) Using midpoints: 5, 15, 25, 35, 45 for the five classes.
Σfx = (4×5) + (10×15) + (18×25) + (12×35) + (6×45) = 20 + 150 + 450 + 420 + 270 = 1310.
Σf = 4+10+18+12+6 = 50.
Mean estimate = 1310 / 50 = 26.2 minutes.

(c) Cumulative frequencies: up to 10 → 4; up to 20 → 14; up to 30 → 32; up to 40 → 44; up to 50 → 50.
The median position is the 25th value (n/2 = 50/2 = 25), which falls in the 20–30 class, since the cumulative frequency reaches 14 by the end of the 10–20 class and 32 by the end of the 20–30 class.
Using linear interpolation: median = 20 + [(25 − 14)/18] × 10 = 20 + (11/18) × 10 = 20 + 6.111... = 26.111... ≈ 26.1 minutes (3 s.f.).

(d) Q1 position = n/4 = 50/4 = 12.5th value, which falls in the 10–20 class (cumulative frequency reaches 4 by the end of 0–10 and 14 by the end of 10–20).
Q1 = 10 + [(12.5 − 4)/10] × 10 = 10 + 8.5 = 18.5 minutes.
Q3 position = 3n/4 = 150/4 = 37.5th value, which falls in the 30–40 class (cumulative frequency reaches 32 by the end of 20–30 and 44 by the end of 30–40).
Q3 = 30 + [(37.5 − 32)/12] × 10 = 30 + (5.5/12) × 10 = 30 + 4.5833... = 34.5833... ≈ 34.6 minutes (3 s.f.).
IQR = Q3 − Q1 = 34.5833 − 18.5 = 16.0833... ≈ 16.1 minutes (3 s.f.).

(e) Quartile coefficient of skewness = (Q3 + Q1 − 2×Median) / (Q3 − Q1)
= (34.5833 + 18.5 − 2×26.1111) / (34.5833 − 18.5)
= (53.0833 − 52.2222) / 16.0833
= 0.8611 / 16.0833 = 0.0535 (3 s.f.).
Since this value is small and positive, the distribution is slightly positively (right) skewed: the gap between the median and Q3 (upper, slower runners) is slightly larger than the gap between Q1 and the median (lower, faster runners), suggesting a small number of runners took noticeably longer than the typical runner, stretching the upper tail of the distribution.
Final answer: (a) 20–30 minutes; (b) mean ≈ 26.2 minutes; (c) median ≈ 26.1 minutes (3 s.f.); (d) Q1 = 18.5 minutes, Q3 ≈ 34.6 minutes (3 s.f.), IQR ≈ 16.1 minutes; (e) quartile skewness coefficient ≈ 0.0535 (3 s.f.), a small positive value indicating the distribution is slightly positively (right) skewed — most runners finished in a fairly typical time, but a longer tail of slower runners pulled the mean and upper quartile further from the median than the faster runners did

评分标准

(a) B1 20–30 minutes. [1]
(b) M1 correct midpoints used; M1 correct Σfx = 1310; M1 divides by Σf = 50; A1 26.2 minutes. [4]
(c) M1 correct cumulative frequencies; M1 correctly identifies median class (20–30) and position (25th); M1 correct linear interpolation formula/substitution; A1 26.1 minutes (accept 26.1 or unrounded 26.11). [4]
(d) M1 correctly finds Q1 = 18.5 (position 12.5, class 10–20); M1 correctly finds Q3 = 34.6 (position 37.5, class 30–40); A1 both Q1 and Q3 correct (18.5 and 34.6 or 34.58); A1 IQR = 16.1 (or 16.08) — follow through from candidate's own Q1, Q3. [4]
(e) M1 correct substitution into the quartile skewness coefficient formula using candidate's own Q1, Q3, median; A1 0.0535 (accept range 0.05 to 0.054 depending on rounding used); A1 correct interpretation (small positive → slight positive/right skew) linked to context (longer tail of slower runners). [3]
Total [16].

想知道自己有几分把握?

thinka 是 DSE 学生在用的 AI 练习应用,提供无限量练习题、即时自动批改和详细解题步骤。超过 100,000 名学生用它确认自己是真的会,而不只是「以为会」。

想练更多同类题型?在 thinka 无限量刷题,即时知道答案。

免费开始练习