CCEA GCSE · thinka 原創模擬試題

2023 CCEA GCSE Statistics 2260 模擬試題連答案詳解

Thinka Jun 2023 CCEA GCSE-Style Mock — Statistics 2260

200 240 分鐘2023
An original Thinka practice paper modelled on the structure and difficulty of the Jun 2023 CCEA GCSE Statistics 2260 paper. Not affiliated with or reproduced from CCEA.

部分 Unit 1 (With Calculator) Higher Tier

Answer all ten questions. Write your answers in the spaces provided. Complete in black ink only. Full working must be clearly shown.
10 題目 · 100
題目 1 · Data Representation & Frequency Polygons
6
1 The table shows the time, t minutes, taken by 40 students to complete a puzzle.
Time (t minutes) | Frequency
0 <= t < 10 | 4
10 <= t < 20 | 9
20 <= t < 30 | 14
30 <= t < 40 | 8
40 <= t < 50 | 5
(a) Write down the modal class. [1]
(b) A frequency polygon is drawn by plotting frequency against the midpoint of each class, and joining the points with straight lines. State the coordinates of the point that would be plotted for the class 30 <= t < 40. [2]
(c) Explain why a frequency polygon, rather than a bar chart, is an appropriate way to represent this continuous grouped data. [3]
查看答案詳解

解題

(a) The class with the highest frequency (14) is 20 <= t < 30, so this is the modal class.
(b) The midpoint of the class 30 <= t < 40 is \( \dfrac{30+40}{2} = 35 \). The frequency for this class is 8, so the point plotted is (35, 8).
(c) Time is continuous data, so there are no natural gaps between classes; a frequency polygon (formed by joining midpoints with straight lines) reflects this continuity better than separate bars, clearly shows the overall shape and trend of the distribution (e.g. where it peaks), and, unlike a bar chart, allows a second frequency polygon to be plotted on the same axes for direct visual comparison between two distributions.

評分準則

(a) B1 for 20<=t<30.
(b) M1 for correct midpoint 35; A1 for coordinate (35,8).
(c) A1 for reference to data being continuous; A1 for reference to showing the shape/trend of the distribution; A1 for reference to ease of comparison with another distribution.
題目 2 · Data Classification & Sampling Methods
8
2 (a) A researcher wants to find out about the shopping habits of people in a large town. She stands outside one supermarket on a Tuesday morning and interviews the first 40 shoppers who agree to speak to her. Name this sampling method, and state ONE limitation of using it to represent the whole town's population. [3]
(b) An alternative method involves dividing the town into 10 geographical areas, randomly selecting 3 of these areas, and then surveying every resident within the 3 selected areas. Name this sampling method. [1]
(c) A third method involves the researcher stopping shoppers (using her own judgement about who to approach) until she has interviewed 20 males and 20 females, matching the town's known gender split. Name this sampling method, and explain why it is not a form of random sampling. [4]
查看答案詳解

解題

(a) This is opportunity (convenience) sampling, since the researcher simply selects whoever is conveniently available at that time and place. A limitation is that the sample is unlikely to be representative of the whole town, since it only includes people who happen to shop at that particular supermarket on a Tuesday morning, excluding people who work at that time, shop elsewhere, or shop on different days.
(b) This is cluster sampling, since the population is divided into geographically-based groups (clusters), a random selection of whole clusters is chosen, and everyone within the selected clusters is surveyed.
(c) This is quota sampling, since the researcher fills fixed quotas (20 males, 20 females) that match known characteristics of the population. It is not a form of random sampling because, within each quota, the researcher chooses which individuals to approach based on her own judgement, so not every member of the population has a known, non-zero, or equal chance of being selected, unlike in true random sampling.

評分準則

(a) B1 for 'opportunity sampling' (accept 'convenience sampling'); B2 for a valid, developed limitation ([1] for a basic/vague limitation).
(b) B1 for 'cluster sampling'.
(c) B1 for 'quota sampling'; B3 for a full, correct explanation of why it is not random ([1]-[2] for a partial explanation, e.g. only stating researcher chooses who to ask without explaining the consequence for selection probability).
題目 3 · Risk & Expected Frequency
7
3 In a large study, the probability that a randomly chosen adult in Town A has a particular allergy is 0.08.
(a) Calculate the expected number of adults with the allergy in a random sample of 250 adults from Town A. [2]
(b) In a separate study of 400 adults from Town B, 48 were found to have the allergy. Calculate the risk (as a percentage) of having the allergy in Town B, based on this sample. [2]
(c) Using your answers to parts (a) and (b), compare the risk of the allergy in the two towns, and comment on whether this proves the allergy is genuinely more common in one town than the other. [3]
查看答案詳解

解題

(a) Expected number \( = 0.08 \times 250 = 20 \) adults.
(b) Risk \( = \dfrac{48}{400} \times 100 = 12\% \).
(c) Town A has a known risk of 8%, while the sample from Town B suggests a risk of 12%, which is higher. However, since Town B's figure of 12% is estimated from just one sample of 400 adults rather than the whole population, this difference could simply be due to natural sampling variation rather than a genuine difference in the underlying risk between the two towns; a larger sample, or repeated samples, or a formal statistical test, would be needed before concluding that the allergy is truly more common in Town B.

評分準則

(a) M1 for 0.08x250; A1 for 20.
(b) M1 for 48/400; A1 for 12%.
(c) A1 for correctly comparing 12% (Town B) with 8% (Town A); A1 for stating this alone does not prove a genuine difference; A1 for a valid reason (e.g. sampling variation, single sample, need for larger sample/test).
題目 4 · Stem-and-Leaf, Boxplots & Distribution Comparison
16
4 The stem-and-leaf diagram shows the time, in minutes, taken by 20 runners (Group A) to complete a fun run. Key: 1 | 5 means 15 minutes.
1 | 5 6 8 9
2 | 0 1 2 3 4 5 6 7 8 9
3 | 0 1 2 3 5 8
(a) Find the median time for Group A. [2]
(b) Find the lower quartile (Q1) and upper quartile (Q3) for Group A, and hence the interquartile range (IQR). [5]
(c) Group B's times (also 20 runners) are summarised by: minimum = 12, Q1 = 18, median = 22, Q3 = 24, maximum = 35 (minutes). Compare the median and the IQR of Group A and Group B, in context. [5]
(d) It is later discovered that the value of 15 minutes recorded for Group A was an error: the runner's true time was 51 minutes. Explain what effect correcting this error would have on (i) the median and (ii) the range of Group A's data, giving the new value in each case. [3]
(e) State ONE reason, other than a recording error, why a genuine outlier might occur in data of this type. [1]
查看答案詳解

解題

(a) The 20 ordered values are: 15,16,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,35,38. The median is the mean of the 10th and 11th values: \( \dfrac{25+26}{2} = 25.5 \) minutes.
(b) Q1 is the median of the lower 10 values (15 to 25): \( \dfrac{20+21}{2} = 20.5 \). Q3 is the median of the upper 10 values (26 to 38): \( \dfrac{30+31}{2} = 30.5 \). IQR \( = 30.5 - 20.5 = 10 \) minutes.
(c) Group A's median (25.5 minutes) is higher than Group B's median (22 minutes), so runners in Group A typically took longer to complete the run than runners in Group B. Group A's IQR (10 minutes) is larger than Group B's IQR (24-18=6 minutes), so the middle 50% of times in Group A are more spread out/variable than in Group B, whose times were more consistent.
(d) (i) Removing 15 and inserting 51 into the sorted data gives: 16,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,35,38,51. The new median is the mean of the 10th and 11th values: \( \dfrac{26+27}{2}=26.5 \) minutes, an increase from 25.5. (ii) The new range is \( 51-16 = 35 \) minutes, compared with the original range of \( 38-15=23 \) minutes, a large increase, since range is very sensitive to extreme values, whereas the median only changed slightly because it depends on the middle values of the ordered data.
(e) A genuine outlier could occur, for example, if one runner was an exceptionally fast, elite-level athlete, or if a runner stopped during the race to help an injured competitor, genuinely increasing their time.

評分準則

(a) M1 for identifying 10th/11th values (25,26); A1 for median 25.5.
(b) M1 for correct method for Q1 (median of lower half); A1 for Q1=20.5; M1 for correct method for Q3 (median of upper half); A1 for Q3=30.5; A1 for IQR=10.
(c) A2 for correct median comparison in context (Group A higher/took longer) ([1] for stating values without context); A2 for correct IQR comparison in context (Group A more variable) ([1] for stating values without context); A1 for correct calculation of Group B's IQR (6).
(d) A1 for new median 26.5 with correct reasoning; A1 for new range 35; A1 for explaining why range changes much more than median (sensitivity to extreme values).
(e) B1 for a valid, plausible reason for a genuine outlier.
題目 5 · Statistical Process Control Charts
11
5 A machine fills bottles with a target mean volume of 500 ml. When the process is in control, the standard deviation of sample means is 4 ml.
(a) Calculate the warning limits (mean +/- 2 standard deviations) and the action limits (mean +/- 3 standard deviations) for a control chart of this process. [4]
(b) The following 8 sample means (ml) are recorded in order: 501, 503, 495, 489, 507, 510, 513, 499. State which sample mean(s), if any, fall beyond the action limits, and explain what the factory should do as a result. [4]
(c) State ONE assumption made when constructing this control chart. [1]
(d) Give ONE consequence to the factory of failing to identify and act on an out-of-control signal promptly. [2]
查看答案詳解

解題

(a) Warning limits \( = 500 \pm 2(4) = 500 \pm 8 \), i.e. 492 ml to 508 ml. Action limits \( = 500 \pm 3(4) = 500 \pm 12 \), i.e. 488 ml to 512 ml.
(b) Checking each value against the action limits [488, 512]: 501, 503, 495, 489, 507 and 499 are all within the action limits; 510 is within the action limits but beyond the upper warning limit (508); 513 ml is beyond the upper action limit (512 ml). Since a sample mean has fallen beyond the action limit, this signals the process is out of control; the factory should stop the production line immediately and investigate/adjust the machine before continuing production.
(c) One assumption is that the process mean (500 ml) and standard deviation (4 ml) used to set the limits are accurate values obtained from a period when the process was known to be running correctly (in control).
(d) If an out-of-control signal is not identified and acted on promptly, the factory may continue producing bottles that are significantly overfilled or underfilled for an extended period, wasting raw materials (if overfilled) or risking under-delivery to customers and regulatory/legal issues (if underfilled), as well as damaging customer trust once the problem is discovered.

評分準則

(a) M1 for correct method for warning limits; A1 for 492 and 508; M1 for correct method for action limits; A1 for 488 and 512.
(b) A1 for correctly identifying 513 ml as beyond the action limit; A1 for correctly stating no other value breaches the action limits; A1 for stating the process is out of control; A1 for a valid recommended action (stop and investigate the process).
(c) B1 for a valid assumption (e.g. limits based on a genuinely in-control historical process, or that data would be normally distributed when in control).
(d) B2 for a valid, developed consequence ([1] for a basic/vague consequence).
題目 6 · Weighted Mean & Algebraic Target Calculations
10
6 (a) A student's coursework mark is 68%, weighted at 30% of the final grade, and their exam mark is x%, weighted at 70% of the final grade. Given that the student's overall weighted mean mark is exactly 75%, form and solve an equation to find x. [5]
(b) A separate small data set of 5 numbers has a mean of 12. Four of the numbers are 8, 10, 14 and 15. Find the fifth number. [3]
(c) State ONE advantage of using the mean, rather than the median, as an average for a data set with no extreme outliers. [2]
查看答案詳解

解題

(a) \( 0.3(68) + 0.7(x) = 75 \Rightarrow 20.4 + 0.7x = 75 \Rightarrow 0.7x = 54.6 \Rightarrow x = 78 \).
(b) Since the mean of 5 numbers is 12, their sum is \( 5 \times 12 = 60 \). The sum of the four known numbers is \( 8+10+14+15=47 \). The fifth number is \( 60-47=13 \).
(c) The mean uses every value in the data set in its calculation, making it a more representative measure of the whole data set (when there are no extreme outliers) than the median, which only depends on the middle value(s); the mean is also more useful for further statistical calculations, such as finding the standard deviation.

評分準則

(a) M1 for correct weighted mean equation set up; M1 for correctly simplifying (20.4+0.7x=75); M1 for correctly isolating 0.7x=54.6; A1 for correct rearrangement; A1 for x=78.
(b) M1 for total = 5x12=60; M1 for sum of four known values = 47; A1 for fifth number = 13.
(c) B2 for a valid, developed advantage ([1] for a basic/vague advantage).
題目 7 · Grouped Frequency Mean, Standard Deviation & Comparison
12
7 The table shows the number of text messages sent in a day by 30 students.
Number of texts (x) | Frequency (f)
5 | 4
10 | 8
15 | 10
20 | 6
25 | 2
(a) Calculate the mean number of texts sent. [3]
(b) Using the formula \( \sqrt{\dfrac{\sum fx^2}{\sum f} - \left(\dfrac{\sum fx}{\sum f}\right)^2} \), calculate the standard deviation of the number of texts sent, giving your answer to 3 significant figures. [5]
(c) A similar survey of 30 adults found a mean of 9 texts per day, with a standard deviation of 3.2. Compare the number of texts sent by the students and by the adults. [4]
查看答案詳解

解題

(a) \( \sum f = 4+8+10+6+2 = 30 \). \( \sum fx = 5(4)+10(8)+15(10)+20(6)+25(2) = 20+80+150+120+50 = 420 \). Mean \( = \dfrac{420}{30} = 14 \).
(b) \( \sum fx^2 = 25(4)+100(8)+225(10)+400(6)+625(2) = 100+800+2250+2400+1250 = 6800 \). Standard deviation \( = \sqrt{\dfrac{6800}{30} - 14^2} = \sqrt{226.67 - 196} = \sqrt{30.67} = 5.54 \) (3 s.f.).
(c) The students' mean (14 texts) is higher than the adults' mean (9 texts), so students send more texts per day on average than adults. The students' standard deviation (5.54) is higher than the adults' standard deviation (3.2), so the number of texts sent by students is more variable (spread out) than the number sent by adults, whose texting habits are more consistent.

評分準則

(a) M1 for Σf=30; M1 for Σfx=420; A1 for mean=14.
(b) M1 for Σfx²=6800; M1 for correct substitution into formula; M1 for 6800/30-196 = 30.67 (or equivalent unsimplified expression); A1 for correct square root method; A1 for SD=5.54 (3 s.f., own-figure rule applies).
(c) A2 for correct mean comparison in context ([1] for stating values without context); A2 for correct SD comparison in context ([1] for stating values without context).
題目 8 · Capture-Recapture Method & Assumptions
7
8 To estimate the population of fish in a lake, a biologist catches and tags 60 fish, then releases them back into the lake. A week later, she catches a second sample of 90 fish, of which 18 are found to be tagged.
(a) Use the capture-recapture (Petersen) method to estimate the total population of fish in the lake. [3]
(b) State TWO assumptions that must hold for this method to give a reliable estimate. [2]
(c) Explain how the population estimate might be affected if some tagged fish had died between the two catches. [2]
查看答案詳解

解題

(a) Using \( N = \dfrac{n_1 \times n_2}{m} \), where \( n_1=60 \) (first catch, tagged), \( n_2=90 \) (second catch) and \( m=18 \) (tagged fish recaptured): \( N = \dfrac{60 \times 90}{18} = \dfrac{5400}{18} = 300 \) fish.
(b) Two assumptions needed are: the population is closed, meaning there are no births, deaths, or migration into or out of the lake between the two catches; and the tagged fish mix back in randomly and evenly with the rest of the population, so that the second catch is representative of the whole lake (all fish equally likely to be caught).
(c) If some tagged fish died before the second catch, fewer tagged fish would be available to be recaptured, so the number of tagged fish found in the second sample (m) would tend to be smaller than it would be if all 60 tagged fish had survived. Since a smaller value of m makes the calculated estimate \( N = n_1n_2/m \) larger, this would cause the method to overestimate the true population of fish in the lake.

評分準則

(a) M1 for correct formula N=n1n2/m; M1 for correct substitution; A1 for N=300.
(b) B1 for each of two valid, distinct assumptions (e.g. closed population; random mixing; tagging doesn't affect survival/catchability; equal catchability), up to [2].
(c) A1 for correctly identifying that m would be smaller than expected; A1 for correctly concluding the estimate would be too high (an overestimate), with valid reasoning linking to the formula.
題目 9 · Spearman's Rank Correlation & Linear Regression Validity
9
9 Six students' rank positions in a Maths test and a Science test are shown below.
Student | A | B | C | D | E | F
Maths rank | 1 | 2 | 3 | 4 | 5 | 6
Science rank | 2 | 1 | 4 | 3 | 6 | 5
(a) Using the formula \( r_s = 1 - \dfrac{6\sum d^2}{n(n^2-1)} \), calculate the value of Spearman's rank correlation coefficient for this data, giving your answer to 3 significant figures. [5]
(b) Interpret, in context, the value you found in part (a). [2]
(c) A teacher says: 'Since rs is close to 1, doing well in Maths causes a student to do well in Science.' Explain why this conclusion is not necessarily valid. [2]
查看答案詳解

解題

(a) The differences in rank, d, are: A: -1, B: 1, C: -1, D: 1, E: -1, F: 1, so \( d^2 = 1 \) for each student, giving \( \sum d^2 = 6 \). With \( n=6 \): \( r_s = 1 - \dfrac{6(6)}{6(6^2-1)} = 1 - \dfrac{36}{6(35)} = 1 - \dfrac{36}{210} = 1 - 0.1714 = 0.829 \) (3 s.f.).
(b) A value of rs = 0.829 indicates a strong positive correlation between students' rank in Maths and their rank in Science: students who rank highly (i.e. do well) in Maths tend to also rank highly in Science, and vice versa.
(c) A strong correlation between two variables does not prove that one causes the other; there could be a third, underlying factor (such as general academic ability, effort, or study habits) that independently causes a student to perform well in both Maths and Science, without either subject directly causing success in the other.

評分準則

(a) M1 for correct differences d; M1 for Σd²=6; M1 for correct substitution into formula; A1 for unsimplified value (1-36/210); A1 for rs=0.829 (3 s.f.).
(b) A2 for a correct, contextualised interpretation referencing strong positive correlation ([1] for a partial/vague interpretation).
(c) A1 for stating correlation does not imply causation; A1 for a valid, specific alternative explanation (e.g. a named plausible third factor).
題目 10 · Normal Distribution Empirical Rule & Standardised z-Scores
14
10 The heights of adult women in a population are normally distributed with mean 165 cm and standard deviation 6 cm.
(a) Using the empirical rule, state the percentage of women with height between 159 cm and 171 cm. [2]
(b) Using the empirical rule, state the percentage of women with height between 153 cm and 177 cm. [2]
(c) Calculate the standardised score (z-score) for a woman of height 174 cm. [3]
The heights of adult men in a different population are normally distributed with mean 178 cm and standard deviation 7 cm.
(d) A man has height 190 cm. Calculate his z-score. [3]
(e) Using standardised scores, determine which is relatively taller compared to their own population: the woman in part (c), or the man in part (d). Give a reason for your choice. [4]
查看答案詳解

解題

(a) 159 cm and 171 cm are exactly 1 standard deviation below and above the mean (165 +/- 6). By the empirical rule, approximately 68% of values in a normal distribution lie within 1 standard deviation of the mean.
(b) 153 cm and 177 cm are exactly 2 standard deviations below and above the mean (165 +/- 12). By the empirical rule, approximately 95% of values lie within 2 standard deviations of the mean.
(c) \( z = \dfrac{174-165}{6} = \dfrac{9}{6} = 1.5 \).
(d) \( z = \dfrac{190-178}{7} = \dfrac{12}{7} = 1.71 \) (3 s.f.).
(e) The man's z-score (1.71) is greater than the woman's z-score (1.5). Since a z-score measures how many standard deviations a value is above (or below) its own population's mean, the man's height is relatively further above his own population's average height than the woman's height is above hers, so the man is relatively taller compared to his own population, even though his actual height in cm and his population's standard deviation both differ from the woman's.

評分準則

(a) B2 for 68% (accept 68.2%).
(b) B2 for 95% (accept 95.4%).
(c) M1 for correct method (174-165)/6; A1 for correct unsimplified value; A1 for z=1.5.
(d) M1 for correct method (190-178)/7; A1 for correct unsimplified value; A1 for z=1.71 (3 s.f.).
(e) A1 for correctly identifying the man as relatively taller; A1 for correctly comparing the two z-scores (1.71 > 1.5); A2 for a correct, fully explained reason referencing what a z-score represents (ft from (c) and (d)).

準備好測試自己了嗎?

將這些筆記轉化為考試練習。獲取此課題的無限量AI題目,即時批改及詳細解析。

練習此課題

部分 Unit 2 (With Calculator) Higher Tier

Answer all ten questions. Write your answers in the spaces provided. Complete in black ink only. Full working must be clearly shown.
10 題目 · 100
題目 1 · Tabular Data, Line Charts & Misleading Visualisations
5
1 A company's advertising claims sales have 'shot up dramatically', showing a bar chart in which the vertical (sales) axis starts at £40,000 rather than £0, with two bars: Year 1 = £42,000 and Year 2 = £48,000.
(a) Calculate the actual percentage increase in sales from Year 1 to Year 2. [2]
(b) Explain how starting the vertical axis at £40,000, rather than £0, makes the increase in sales appear misleadingly large. [3]
查看答案詳解

解題

(a) Percentage increase \( = \dfrac{48000-42000}{42000} \times 100 = \dfrac{6000}{42000}\times100 = 14.3\% \) (3 s.f.).
(b) By starting the vertical axis at £40,000 instead of £0, the visible portion of each bar only represents sales above £40,000 (£2,000 for Year 1 and £8,000 for Year 2), so the second bar appears four times taller than the first, even though the underlying increase in total sales is a much more modest 14.3%. This truncated axis exaggerates the visual impression of growth, making a moderate increase look dramatic.

評分準則

(a) M1 for correct method (6000/42000); A1 for 14.3%.
(b) A1 for identifying that only the portion above £40,000 is shown/compared; A1 for correctly explaining this exaggerates the relative visual difference between the bars; A1 for linking this back to the true, much smaller, percentage increase found in (a).
題目 2 · Sampling Frame, Hypotheses & Questionnaire Design
9
2 A student wants to test the hypothesis: 'Students who walk to school get more sleep than students who are driven to school.'
(a) Write a suitable null hypothesis for this investigation. [2]
(b) Suggest an appropriate sampling frame the student could use to select a sample of students from their school. [2]
(c) The student designs this questionnaire item: 'Don't you agree that walking to school is much healthier than being driven?' Identify TWO problems with this question, and rewrite it as a more suitable, unbiased pair of questions that would let the student test the hypothesis. [5]
查看答案詳解

解題

(a) A suitable null hypothesis is: 'There is no difference in the average amount of sleep obtained by students who walk to school and students who are driven to school' (i.e. mode of travel to school and amount of sleep are independent/unrelated).
(b) An appropriate sampling frame would be the school's full register (list) of all currently enrolled students, from which a sample could then be randomly selected.
(c) Two problems with the question are: it is a leading question ('Don't you agree...'), which pressures respondents towards answering 'yes'; and it does not actually collect the factual data needed to test the hypothesis (travel mode and amount of sleep), instead asking for an opinion on health that is not directly relevant. A more suitable pair of questions would be, for example: 'How do you usually travel to school? (tick one): Walk / Driven / Other' and 'On an average school night, how many hours of sleep do you get? _____ hours', which are neutral, closed, and collect exactly the data needed.

評分準則

(a) B2 for a correctly worded null hypothesis referencing no difference/independence between travel mode and sleep ([1] for a partial/vague statement).
(b) B2 for a valid sampling frame (e.g. school register/list of all students) ([1] for a vague/partial answer).
(c) B1 for identifying the question is leading/biased; B1 for identifying the question does not collect the relevant factual data; B3 for a genuinely improved, closed, unbiased rewritten question (or pair of questions) that would allow the hypothesis to be tested ([1]-[2] for a partial improvement).
題目 3 · Quality Control Limits & Operational Evaluation
14
3 A packing process fills bags of rice with a target mean mass of 1000 g. From 50 historical samples taken when the process was known to be in control, the mean of the sample means was 1000 g, with a standard deviation of the sample means of 3 g.
(a) State the warning limits (mean +/- 2 SD) and the action limits (mean +/- 3 SD) for a control chart of this process. [4]
(b) Six new sample means (g) are recorded, in order: 999, 1004, 1007, 996, 988, 1010. State which sample mean(s), if any, fall beyond the action limits, and state what the operator should do as a result. [5]
(c) A further six sample means are then recorded: 1001, 1002, 1001, 1003, 1002, 1001. All lie comfortably within the control limits. Comment on what this unusually tight run of values might suggest about the process or the monitoring equipment. [3]
(d) Give ONE cost to the company of stopping the production line unnecessarily due to a false alarm. [2]
查看答案詳解

解題

(a) Warning limits \( = 1000 \pm 2(3) = 1000 \pm 6 \), i.e. 994 g to 1006 g. Action limits \( = 1000 \pm 3(3) = 1000 \pm 9 \), i.e. 991 g to 1009 g.
(b) Checking each value against the action limits [991,1009]: 999, 1004, 1007 and 996 are within the action limits (1007 is beyond the warning limit of 1006 but not the action limit); 988 g is below the lower action limit (991); 1010 g is above the upper action limit (1009). Both 988 g and 1010 g fall beyond the action limits, signalling the process is out of control; the operator should stop the production line and investigate/adjust the machine.
(c) A run of six sample means clustered this tightly around 1001-1003 g, with far less variation than the historical standard deviation of 3 g would suggest, could indicate a genuine improvement in the consistency of the filling process. However, it could equally suggest a problem with the monitoring equipment itself, such as a stuck or faulty weighing gauge that is repeatedly giving very similar readings rather than accurately measuring each sample, so this pattern should prompt a check of the measuring equipment as well as the filling process.
(d) One cost of stopping the line unnecessarily (a 'false alarm', where the process was actually still in control) is the lost production output and staff time while the line is halted and checked, reducing the factory's overall productivity and potentially causing delays in meeting customer orders.

評分準則

(a) M1 for correct method for warning limits; A1 for 994 and 1006; M1 for correct method for action limits; A1 for 991 and 1009.
(b) A1 for correctly identifying 988g beyond the action limit; A1 for correctly identifying 1010g beyond the action limit; A1 for correctly stating no other values breach the action limits; A1 for stating the process is out of control; A1 for a valid recommended action.
(c) A1 for suggesting genuine improved consistency as one possibility; A1 for suggesting a fault with the measuring equipment as a second possibility; A1 for a valid recommended follow-up action (checking the equipment).
(d) B2 for a valid, developed cost of an unnecessary stoppage ([1] for a basic/vague cost).
題目 4 · Time Series Graphs, Trends & PMCC Calculation
10
4 The table shows the advertising spend, x (£000s), and the sales, y (£000s), for 5 shops in a chain.
Shop | 1 | 2 | 3 | 4 | 5
Advertising spend x | 2 | 4 | 6 | 8 | 10
Sales y | 20 | 28 | 35 | 45 | 52
[You are given: n=5, sum(x)=30, sum(y)=180, sum(xy)=1242, sum(x^2)=220, sum(y^2)=7138.]
(a) Calculate the product moment correlation coefficient (PMCC) for this data, giving your answer to 3 significant figures. [5]
(b) Interpret this value in context. [2]
(c) A different shop's sales over the past 5 years, shown on a time series line graph, rose steadily from £20,000 to £35,000 over the first 3 years, then fell to £30,000 in Year 4, before rising sharply to £42,000 in Year 5. Describe the trend shown by this time series. [3]
查看答案詳解

解題

(a) \( r = \dfrac{n\sum xy - \sum x \sum y}{\sqrt{\left(n\sum x^2-(\sum x)^2\right)\left(n\sum y^2-(\sum y)^2\right)}} = \dfrac{5(1242)-(30)(180)}{\sqrt{\left(5(220)-30^2\right)\left(5(7138)-180^2\right)}} \)
\( = \dfrac{6210-5400}{\sqrt{(1100-900)(35690-32400)}} = \dfrac{810}{\sqrt{200 \times 3290}} = \dfrac{810}{\sqrt{658000}} = \dfrac{810}{811.17} = 0.999 \) (3 s.f.).
(b) A PMCC of 0.999 indicates a very strong positive (linear) correlation between advertising spend and sales for these shops: as advertising spend increases, sales increase in an almost perfectly linear way.
(c) Overall, the time series shows an upward trend in sales over the 5 years, rising from £20,000 to £42,000. However, this was not a steady rise throughout: sales rose steadily for the first 3 years, then experienced a temporary fall in Year 4 (from £35,000 to £30,000), before recovering sharply and rising to a new high of £42,000 by Year 5.

評分準則

(a) M1 for correct numerator (6210-5400=810); M1 for correct first bracket (1100-900=200); M1 for correct second bracket (35690-32400=3290); A1 for correct unsimplified expression; A1 for r=0.999 (3 s.f.).
(b) A2 for a correct, contextualised interpretation referencing very strong positive correlation ([1] for a partial/vague interpretation).
(c) A1 for correctly identifying the overall upward trend; A1 for correctly identifying the Year 4 dip; A1 for correctly identifying the sharp Year 5 recovery.
題目 5 · Two-Way Tables, Frequency Trees & Conditional Probability
13
5 Of 200 students at a school, 90 study French, 70 study Spanish, and 30 study both French and Spanish.
(a) Complete a two-way table showing the number of students who do/do not study French, against the number who do/do not study Spanish (state all four missing cell values, given the totals above). [4]
(b) A student is selected at random. Find the probability that they study French but not Spanish. [2]
(c) A student is selected at random from those who study Spanish. Find the probability that they also study French, i.e. P(French | Spanish). [2]
(d) Describe a frequency tree showing the split of the 200 students first by French/not French, then by Spanish/not Spanish within each branch, and use it to find the number of students who study neither language. [3]
(e) State, showing your working, whether studying French and studying Spanish are independent events. [2]
查看答案詳解

解題

(a) Since 90 study French and 30 of these also study Spanish, 60 study French but not Spanish. Since 70 study Spanish and 30 of these also study French, 40 study Spanish but not French. The remaining students study neither: \( 200-(30+60+40)=70 \).
Two-way table:
| Studies Spanish | Doesn't study Spanish | Total
Studies French | 30 | 60 | 90
Doesn't study French | 40 | 70 | 110
Total | 70 | 130 | 200
(b) \( P(\text{French not Spanish}) = \dfrac{60}{200} = 0.3 \).
(c) \( P(\text{French}|\text{Spanish}) = \dfrac{30}{70} = \dfrac{3}{7} \approx 0.429 \).
(d) The frequency tree starts with 200 students, splitting into 90 (French) and 110 (Not French). The French branch (90) splits into 30 (Spanish) and 60 (Not Spanish). The Not French branch (110) splits into 40 (Spanish) and 70 (Not Spanish). The number of students studying neither language is the 'Not French, Not Spanish' branch: 70 students.
(e) \( P(\text{French}) = \dfrac{90}{200} = 0.45 \). \( P(\text{French}|\text{Spanish}) = \dfrac{30}{70} \approx 0.429 \). Since \( P(\text{French}|\text{Spanish}) \neq P(\text{French}) \) (0.429 is not equal to 0.45), studying French and studying Spanish are not independent events.

評分準則

(a) B1 for each correct missing cell value (30 given/not required to derive; 60, 40, 70), up to [4] (accept correctly completed table in any equivalent layout).
(b) M1 for 60/200; A1 for 0.3.
(c) M1 for 30/70; A1 for 3/7 or 0.429 (accept equivalent).
(d) B1 for correct first-level split (90/110); B1 for correct second-level splits (30/60 and 40/70); B1 for correctly identifying neither=70.
(e) M1 for correctly calculating both P(French) and P(French|Spanish); A1 for correct conclusion (not independent) with valid comparison of the two values.
題目 6 · Venn Diagrams & Intersection/Union Probability
8
6 In a group of 60 students, 35 like football, 28 like basketball, and 12 like both sports.
(a) Describe the numbers that would appear in each region of a Venn diagram for football (F) and basketball (B), including the number of students who like neither sport. [4]
(b) Find P(F union B), the probability that a randomly chosen student likes at least one of the two sports. [2]
(c) Find P(football only). [2]
查看答案詳解

解題

(a) The intersection (both sports) region contains 12 students. The 'football only' region contains \( 35-12=23 \) students. The 'basketball only' region contains \( 28-12=16 \) students. The number of students represented within the two circles is \( 23+16+12=51 \), so the region outside both circles (neither sport) contains \( 60-51=9 \) students.
(b) \( P(F \cup B) = \dfrac{51}{60} = 0.85 \) (or equivalently \( \dfrac{17}{20} \)).
(c) \( P(\text{football only}) = \dfrac{23}{60} \approx 0.383 \) (3 s.f.).

評分準則

(a) B1 for both=12; B1 for football only=23; B1 for basketball only=16; B1 for neither=9.
(b) M1 for 51/60 (ft from (a)); A1 for 0.85 (or 17/20).
(c) M1 for 23/60 (ft from (a)); A1 for 0.383 (accept equivalent fraction/decimal).
題目 7 · Cumulative Frequency Diagrams, Quartiles & Boxplot Comparison
15
7 The table shows the distance travelled to school, d km, by 80 students at an urban school.
Distance (d km) | Frequency | Cumulative frequency
0 <= d < 2 | 12 | ?
2 <= d < 4 | 20 | ?
4 <= d < 6 | 24 | ?
6 <= d < 8 | 16 | ?
8 <= d < 10 | 8 | ?
(a) Complete the cumulative frequency column. [2]
(b) Using linear interpolation on the cumulative frequency table, estimate the median distance travelled. [3]
(c) Estimate the lower quartile (Q1) and upper quartile (Q3), and hence the interquartile range (IQR). [5]
A similar survey of 80 students at a rural school gives the five-number summary: minimum = 1, Q1 = 5, median = 8, Q3 = 12, maximum = 20 (km).
(d) Compare the distances travelled by students at the urban school (parts (a)-(c)) and the rural school, in context. [4]
(e) State ONE reason why students at the rural school might travel further to school than students at the urban school. [1]
查看答案詳解

解題

(a) Cumulative frequencies: 12, 32, 56, 72, 80.
(b) The median is the \( \tfrac{80}{2}=40\text{th} \) value, which falls in the class 4<=d<6 (cumulative frequency rises from 32 to 56 in this class). Using linear interpolation: \( \text{median} = 4 + \dfrac{40-32}{56-32}\times 2 = 4 + \dfrac{8}{24}\times2 = 4.67 \) km (3 s.f.).
(c) Q1 is the \( \tfrac{80}{4}=20\text{th} \) value, which falls in the class 2<=d<4: \( Q_1 = 2 + \dfrac{20-12}{32-12}\times2 = 2+\dfrac{8}{20}\times2 = 2.8 \) km. Q3 is the \( \tfrac{3(80)}{4}=60\text{th} \) value, which falls in the class 6<=d<8: \( Q_3 = 6+\dfrac{60-56}{72-56}\times2 = 6+\dfrac{4}{16}\times2=6.5 \) km. IQR \( = 6.5-2.8=3.7 \) km.
(d) The rural school's median (8 km) is higher than the urban school's median (4.67 km), so rural students typically travel further to school than urban students. The rural school's IQR (12-5=7 km) is also higher than the urban school's IQR (3.7 km), so the distances travelled by rural students are more variable/spread out than those travelled by urban students, whose journeys are more consistent.
(e) Rural areas typically have a lower population density than urban areas, meaning homes and the nearest school are, on average, further apart, and there may be fewer schools to choose from locally.

評分準則

(a) B2 for all correct cumulative frequencies (12,32,56,72,80) ([1] for at least 3 correct).
(b) M1 for correctly identifying the 40th value falls in class 4<=d<6; M1 for correct interpolation method; A1 for median=4.67 km (accept 4.66-4.67).
(c) M1 for correctly identifying the 20th value's class; A1 for Q1=2.8 km; M1 for correctly identifying the 60th value's class; A1 for Q3=6.5 km; A1 for IQR=3.7 km.
(d) A2 for correct median comparison in context (rural higher/travel further) ([1] for stating values without context); A2 for correct IQR comparison in context (rural more variable) ([1] for stating values without context, and for correctly calculating rural IQR=7).
(e) B1 for a valid, developed reason (e.g. lower population density, schools further apart, fewer local schools).
題目 8 · Binomial Distribution Probability Expansion
8
8 A multiple-choice quiz has 5 questions, each with 4 options, only one of which is correct. A student guesses the answer to every question independently at random.
(a) State TWO conditions needed for the number of correct guesses to be modelled by a binomial distribution. [2]
(b) Using X ~ B(5, 0.25), calculate P(X = 2), the probability the student gets exactly 2 questions correct. Give your answer to 3 significant figures. [4]
(c) Calculate P(X >= 1), the probability the student gets at least one question correct. [2]
查看答案詳解

解題

(a) Two conditions needed are: there is a fixed number of trials (here, 5 questions); each trial is independent of the others; each trial has only two possible outcomes (correct or incorrect); and the probability of success (correct guess, 0.25) is constant across all trials.
(b) \( P(X=2) = \binom{5}{2}(0.25)^2(0.75)^3 = 10 \times 0.0625 \times 0.421875 = 0.264 \) (3 s.f.).
(c) \( P(X \geq 1) = 1-P(X=0) = 1-(0.75)^5 = 1-0.237 = 0.763 \) (3 s.f.).

評分準則

(a) B1 for each of two valid, distinct conditions, up to [2] (e.g. fixed number of trials; independence; only two outcomes; constant probability).
(b) M1 for correct binomial coefficient C(5,2)=10; M1 for correct powers 0.25² and 0.75³; A1 for correct unsimplified expression; A1 for P(X=2)=0.264 (3 s.f.).
(c) M1 for correct method 1-(0.75)^5; A1 for P(X>=1)=0.763 (3 s.f.).
題目 9 · Chain Base Index Numbers & Percentage Interpretations
8
9 The table shows the chain base price index for a raw material over 4 years (each year's index is calculated relative to the PREVIOUS year, which is always given a value of 100).
Year | Chain base index
2020 | 100 (base)
2021 | 105
2022 | 98
2023 | 110
(a) Interpret, in context, the chain base index value for 2021. [2]
(b) Interpret, in context, the chain base index value for 2022. [2]
(c) Given that the actual price in 2020 was £40 per unit, calculate the actual price in 2021. [2]
(d) Continuing the chain from your answer to part (c), calculate the overall percentage change in price from 2020 to 2023. [2]
查看答案詳解

解題

(a) A chain base index of 105 for 2021 means the price rose by 5% between 2020 and 2021.
(b) A chain base index of 98 for 2022 means the price fell by 2% between 2021 and 2022.
(c) Price in 2021 \( = £40 \times 1.05 = £42 \).
(d) Continuing the chain: price in 2022 \( = £42 \times 0.98 = £41.16 \); price in 2023 \( = £41.16 \times 1.10 = £45.28 \) (2 d.p.). Overall percentage change from 2020 to 2023: \( \dfrac{45.28-40}{40}\times100 = 13.2\% \) (3 s.f.) increase.

評分準則

(a) A2 for correct interpretation (price rose by 5%, 2020 to 2021) ([1] for a partial/vague interpretation).
(b) A2 for correct interpretation (price fell by 2%, 2021 to 2022) ([1] for a partial/vague interpretation).
(c) M1 for 40x1.05; A1 for £42 (ft from correct method even if (a) misread).
(d) M1 for correctly chaining through 2022 (£41.16) and 2023 (£45.28); A1 for overall percentage change 13.2% (own-figure rule applies throughout).
題目 10 · Sample Standard Deviation & Standardised Score Comparison
10
10 The scores of 8 students in a test are: 12, 15, 18, 20, 14, 16, 19, 22.
(a) Calculate the mean of this data. [2]
(b) Using the formula \( \sqrt{\dfrac{\sum x^2}{n} - \left(\dfrac{\sum x}{n}\right)^2} \), calculate the standard deviation of this data, giving your answer to 3 significant figures. [4]
A student scored 21 in this test. The same student also scored 78 in a different test, where the class mean was 72 and the standard deviation was 5.
(c) Calculate the z-score for each of the student's two test results. [3]
(d) Using your answers to part (c), state in which test the student performed relatively better, giving a reason. [1]
查看答案詳解

解題

(a) \( \sum x = 12+15+18+20+14+16+19+22 = 136 \). Mean \( = \dfrac{136}{8} = 17 \).
(b) \( \sum x^2 = 144+225+324+400+196+256+361+484 = 2390 \). Standard deviation \( = \sqrt{\dfrac{2390}{8} - 17^2} = \sqrt{298.75-289} = \sqrt{9.75} = 3.12 \) (3 s.f.).
(c) For the first test: \( z = \dfrac{21-17}{3.12} = 1.28 \) (3 s.f.). For the second test: \( z = \dfrac{78-72}{5} = 1.2 \).
(d) The student's z-score for the first test (1.28) is higher than for the second test (1.2), meaning the score of 21 was relatively further above its class mean (in standard deviation terms) than the score of 78 was above its own class mean. The student therefore performed relatively better, compared with their respective class, in the first test.

評分準則

(a) M1 for Σx=136; A1 for mean=17.
(b) M1 for Σx²=2390; M1 for correct substitution into formula; A1 for 9.75 (unsimplified value under root); A1 for SD=3.12 (3 s.f., own-figure rule applies).
(c) M1 for correct method for both z-scores; A1 for z(first test)=1.28 (ft from (a),(b)); A1 for z(second test)=1.2.
(d) A1 for correctly identifying the first test as relatively better, with valid reasoning comparing the two z-scores (ft from (c)).

想知道自己有幾分把握?

thinka 是 DSE 學生用的 AI 練習應用程式,有無限量練習題、即時自動批改和詳細解題步驟。逾 100,000 名學生用它確認自己真的識,而不只是「以為識」。

想練更多類似題型?在 thinka 無限量操練,即時知道答案。

免費開始練習