Hard A-Level Sampling Questions

Challenging, exam-style A-Level Sampling questions with worked solutions. Stretch yourself on the hardest stratified, rounding, two strata, simple random problems.

stratifiedroundingtwo stratasimple randombiassystematic
A-Level34 questionsStep-by-step solutions
Question 1
9 markschallenging
A researcher plans a national study using a complete electoral register and a computer, and needs high representativeness across regions and age groups. Compare the methods and choose the best, giving reasons.
Show worked solution

Worked solution

  1. State the requirements

    National, representative, register, computer\text{National, representative, register, computer}

    The study is national, needs high representativeness across regions and ages, and has a full register plus a computer. The method must use these strengths.

  2. List the candidate methods

    SRS, systematic, stratified, quota, opportunity\text{SRS, systematic, stratified, quota, opportunity}

    The five standard methods are simple random, systematic, stratified, quota and opportunity. We assess each against the requirements.

  3. Assess opportunity sampling

    Not representative\text{Not representative}

    Opportunity sampling uses whoever is available and ignores regions and ages. So it fails the representativeness requirement.

  4. Assess quota sampling

    Non-random, wastes register\text{Non-random, wastes register}

    Quota sampling controls groups but selects non-randomly and does not use the register. It risks interviewer bias, so it is not ideal.

  5. Assess simple random sampling

    Random but groups by chance\text{Random but groups by chance}

    Simple random sampling uses the register fairly, but region and age balance is left to chance. Small regions could be under-represented.

  6. Assess systematic sampling

    Even spacing, no group control\text{Even spacing, no group control}

    Systematic sampling is easy on a computer but does not guarantee proportional regions and ages. It also risks periodicity in an ordered list.

  7. Assess stratified sampling

    Groups + proportion + random\text{Groups + proportion + random}

    Stratified sampling divides the register into region and age strata and samples each in proportion at random. This directly targets representativeness.

  8. Use the register

    Register gives every member\text{Register gives every member}

    The complete register lets us identify each person's region and age. So we can form the strata accurately.

  9. Use the computer

    Computer allocates and selects\text{Computer allocates and selects}

    A computer can quickly compute proportional shares and pick random members from each stratum. So the extra work of stratification is easy here.

  10. Note the representativeness gain

    Guarantees proportional groups\text{Guarantees proportional groups}

    Stratified sampling guarantees each region and age group appears in proportion. This is exactly what the study needs.

  11. Note the randomness benefit

    Random within strata\text{Random within strata}

    Selection within each stratum is random, so bias is minimised. This beats quota sampling's non-random choice.

  12. Compare with simple random

    Stratified>SRS for balance\text{Stratified} > \text{SRS for balance}

    Unlike simple random sampling, stratified sampling removes the risk of chance imbalance between groups. So it is more representative.

  13. Weigh the trade-off

    Needs group sizes, which we have\text{Needs group sizes, which we have}

    Stratified sampling needs known group sizes, and the register provides them. So the only drawback is covered.

  14. Eliminate the weaker methods

    Others fail a requirement\text{Others fail a requirement}

    Opportunity, quota, simple random and systematic each fail on representativeness or use of the register. So they are eliminated.

  15. Select stratified sampling

    Best fit for all requirements\text{Best fit for all requirements}

    Stratified random sampling meets every requirement using the register and computer. So it is the best method.

  16. Conclude

    Stratified random sampling\text{Stratified random sampling}

    So stratified random sampling is best: the computer allocates proportional random samples to each region and age stratum from the register, giving high representativeness. That is the correct choice.

Answer
Stratified random sampling, because the computer can allocate proportional random samples to each region and age stratum from the register, giving high representativeness.
Question 2
9 markschallenging
A stratified sample equal to 5%5\% of a population of 48004800 is spread over 44 sites of 15001500, 12001200, 11001100 and 10001000. Find the number sampled from the largest site.
Show worked solution

Worked solution

  1. Write the population

    N=4800N = 4800

    The whole population is 4800 people. We first need the total sample size.

  2. Write the sampling fraction

    f=5%=0.05f = 5\% = 0.05

    The sample is 5% of the population, so the sampling fraction is 0.05. This applies to each site.

  3. Find the total sample size

    n=0.05×4800=240n = 0.05 \times 4800 = 240

    Multiplying 4800 by 0.05 gives 240. So the whole sample has 240 people.

  4. Check the sites total

    1500+1200+1100+1000=48001500+1200+1100+1000 = 4800

    Adding the four sites gives 4800, the whole population. So the strata are complete.

  5. Identify the largest site

    max=1500\max = 1500

    The largest site has 1500 people. This is the stratum we are asked about.

  6. Apply the fraction to it

    1500×0.051500 \times 0.05

    We take 5% of the largest site. Multiply 1500 by 0.05.

  7. Evaluate

    1500×0.05=751500 \times 0.05 = 75

    This gives 75 people from the largest site. That is our candidate answer.

  8. Check it is a whole number

    75Z75 \in \mathbb{Z}

    75 is already whole, so no rounding is needed. The value stands.

  9. Compare the population share

    15004800=0.3125\frac{1500}{4800} = 0.3125

    The largest site is about 31.25% of the population. We compare with its sample share.

  10. Compare the sample share

    75240=0.3125\frac{75}{240} = 0.3125

    75 out of 240 is also 31.25%, matching the population share. So the answer is consistent.

  11. Sample the other sites

    1200,1100,1000×0.05=60,55,501200,1100,1000 \times 0.05 = 60,55,50

    For a check, 5% of the other sites gives 60, 55 and 50. These help confirm the total.

  12. Add all shares

    75+60+55+50=24075+60+55+50 = 240

    The four site shares add to exactly 240, the sample size. So the allocation is valid.

  13. Confirm the total

    240=240240 = 240 \checkmark

    The total equals the required sample size, so no adjustment is needed. The largest site's value is confirmed.

  14. Restate the calculation

    0.05×1500=750.05 \times 1500 = 75

    So 5% of the largest site of 1500 is 75. This is the sample from that site.

  15. State the answer

    7575

    So 75 are sampled from the largest site. That is the required value.

Answer
7575
Question 3
9 markschallenging
To find out how the local community uses the public library, a survey is carried out on people \textbf{inside} the library. Identify the flaw and the best fix.
Show worked solution

Worked solution

  1. State the aim

    Study the whole community’s library use\text{Study the whole community's library use}

    The survey is about how the community uses the library, including those who do not. So the sample should cover everyone.

  2. Identify the method

    Survey people inside the library\text{Survey people inside the library}

    Only people currently in the library are asked. This is where the bias arises.

  3. Note who is sampled

    Only current users\text{Only current users}

    Everyone asked is already using the library. So the sample is made up entirely of users.

  4. Identify the excluded group

    Non-users missing\text{Non-users missing}

    People who never use the library are not there to be asked. They are completely excluded.

  5. Link to the aim

    Overstates library use\text{Overstates library use}

    Because only users are surveyed, the results will overstate community use. Non-use is invisible to this method.

  6. Name the flaw

    Coverage bias\text{Coverage bias}

    The sampling location excludes part of the population. This is a coverage or selection bias.

  7. Consider surveying more users

    More usersfix\text{More users} \neq \text{fix}

    Asking more people inside the library still ignores non-users. So volume does not solve it.

  8. Consider a different time

    Different time, still users\text{Different time, still users}

    Surveying at another time inside the library still only reaches users. The exclusion remains.

  9. State what a fix needs

    Reach users and non-users\text{Reach users and non-users}

    A good fix must include people who do not use the library. Everyone in the community should have a chance.

  10. Introduce a community frame

    Use an address/electoral list\text{Use an address/electoral list}

    We should sample from a list of the whole community, such as an address register. This includes non-users.

  11. Choose a fair method

    Random sample of residents\text{Random sample of residents}

    From that frame take a random sample of residents. This gives non-users a chance to appear.

  12. Note the benefit

    Captures non-use too\text{Captures non-use too}

    Now the survey can measure both use and non-use. So it reflects the whole community.

  13. Compare with the flawed method

    Community framein-library\text{Community frame} \gg \text{in-library}

    Sampling the community beats sampling inside the library for representativeness. So this is the right fix.

  14. Eliminate weak redesigns

    Reject in-library-only options\text{Reject in-library-only options}

    Any redesign that still surveys only people in the library keeps the bias. So those options are eliminated.

  15. Conclude

    Sample the whole community at random\text{Sample the whole community at random}

    So the flaw is that only current users are surveyed; the fix is to take a random sample from the whole community, not just people in the library. That is the best answer.

Answer
Everyone surveyed already uses the library, so non-users are excluded; instead survey a random sample of the whole community.
Question 4
9 markschallenging
A simple random sample of 33 is taken from 1010 people. Find the probability that two particular people are both in the sample.
Show worked solution

Worked solution

  1. Set up the problem

    N=10, n=3N = 10,\ n = 3

    We choose 3 people at random from 10, with every group of 3 equally likely. We want two named people both included.

  2. Count all possible samples

    (103)\binom{10}{3}

    The number of possible samples is the number of ways to choose 3 from 10. This is the denominator.

  3. Evaluate the total

    (103)=120\binom{10}{3} = 120

    Working out 10 choose 3 gives 120. So there are 120 equally likely samples.

  4. Set up the favourable count

    Both A and B in, choose 1 more\text{Both A and B in, choose 1 more}

    If the two named people are both in, we only need to choose the remaining member. That comes from the other 8 people.

  5. Count the favourable samples

    (81)=8\binom{8}{1} = 8

    Choosing 1 from the remaining 8 gives 8 ways. So 8 of the samples contain both named people.

  6. Write the probability

    P=8120P = \frac{8}{120}

    The probability is favourable over total, so 8 over 120. Now we simplify.

  7. Simplify

    8120=115\frac{8}{120} = \frac{1}{15}

    Dividing top and bottom by 8 gives one fifteenth. That is the probability.

  8. Set up a check by multiplication

    P=P(A)×P(BA)P = P(A)\times P(B\mid A)

    We can verify using the chance A is in, times the chance B is in given A is. This is an independent route to the answer.

  9. Probability A is chosen

    P(A)=310P(A) = \frac{3}{10}

    Person A has a 3 in 10 chance of being in the sample. This is the sample size over the population.

  10. Probability B given A

    P(BA)=29P(B\mid A) = \frac{2}{9}

    Given A is in, 2 of the remaining 9 places could be B, from 9 people left. So the chance is 2 over 9.

  11. Multiply

    310×29=690\frac{3}{10}\times\frac{2}{9} = \frac{6}{90}

    Multiplying the two probabilities gives 6 over 90. Now simplify.

  12. Simplify the check

    690=115\frac{6}{90} = \frac{1}{15}

    This also reduces to one fifteenth. So both methods agree.

  13. Compare the methods

    Both give 115\text{Both give } \tfrac{1}{15}

    The counting method and the multiplication method give the same answer. This confirms the result.

  14. Interpret

    P=115P = \tfrac{1}{15}

    So there is a 1 in 15 chance that both named people are in the sample. This is small, as expected for a small sample.

  15. State the answer

    P=115P = \frac{1}{15}

    So the required probability is one fifteenth. That is the final answer.

Answer
115\frac{1}{15}
Question 5
9 markschallenging
To study how much exercise people in a town do, a researcher surveys members leaving a gym. Identify the flaw and choose the best redesign.
Show worked solution

Worked solution

  1. State the aim

    Measure exercise across the town\text{Measure exercise across the town}

    The study is about exercise levels in the whole town. So the sample should represent all residents.

  2. Identify the method

    Survey gym-leavers\text{Survey gym-leavers}

    The researcher only asks people leaving a gym. This is where the problem starts.

  3. Note who is sampled

    Only gym members\text{Only gym members}

    Everyone asked already exercises at a gym. So the sample is a special, active group.

  4. Identify the excluded group

    Non-gym people missing\text{Non-gym people missing}

    People who exercise little or not at all rarely visit a gym. They are excluded from the sample.

  5. Link to the aim

    Overstates exercise\text{Overstates exercise}

    Because only active people are asked, the study will overstate the town's exercise. This is a clear bias.

  6. Name the flaw

    Unrepresentative frame\text{Unrepresentative frame}

    The sampling location excludes a key part of the population. So the sample cannot represent the town.

  7. State the goal of a fix

    Include all activity levels\text{Include all activity levels}

    A good redesign must reach both active and inactive people. Everyone in the town should have a chance.

  8. Consider a bigger gym sample

    More gym-goersfix\text{More gym-goers} \neq \text{fix}

    Surveying more gym members does not help, since non-members are still excluded. Volume cannot cure this bias.

  9. Consider a second gym

    Still only gym-goers\text{Still only gym-goers}

    Adding another gym still samples only active people. So the same exclusion remains.

  10. Introduce a proper frame

    Use a town-wide list\text{Use a town-wide list}

    Instead we should sample from a list of the whole town, such as an address or electoral register. This includes everyone.

  11. Choose a fair method

    Random or stratified sample\text{Random or stratified sample}

    From that frame we take a random or stratified sample of residents. This gives all activity levels a chance.

  12. Note the benefit

    Represents inactive people too\text{Represents inactive people too}

    Now inactive people can appear in the sample. So the estimate reflects the whole town.

  13. Compare with the flawed method

    Town framegym exit\text{Town frame} \gg \text{gym exit}

    Sampling the town beats sampling a gym exit for representativeness. So this is the right fix.

  14. Eliminate weak redesigns

    Reject gym-only options\text{Reject gym-only options}

    Any redesign that still only asks gym-goers keeps the bias. So those options are eliminated.

  15. Conclude

    Sample the whole town at random\text{Sample the whole town at random}

    So the flaw is that only gym members are surveyed; the fix is to take a random or stratified sample from the whole town's population. That is the best redesign.

Answer
Everyone surveyed already exercises, so non-exercisers are excluded; instead take a random (or stratified) sample from the whole town's population register.

Unlock 29 more Sampling questions

Create a free account to work through every A-Level Sampling question with instant step-by-step worked solutions, progress tracking and interactive lessons.

  • Full worked solutions for every question
  • Interactive lessons and instant feedback
  • Track your mastery across every topic
Create a Free Account

No card required · Free forever

More Sampling practice

Related Statistics topics