Hard Further Maths Estimation and confidence intervals Questions
Challenging, exam-style Further Maths Estimation and confidence intervals questions with worked solutions. Stretch yourself on the hardest confidence-interval, unknown-variance, t-distribution, required-width problems.
A random sample X1,X2,…,Xn of size n=25 is taken from a population of nylon threads whose breaking strain has unknown mean μ and unknown variance σ2. Which of the following statements concerning unbiased estimators is correct?
Show worked solution
Worked solution
State the definition of an unbiased estimator
E(θ^)=θ
Unbiasedness is a property of the ESTIMATOR across all samples, not of the estimate from one sample.
Take the expectation of the sample mean
E(Xˉ)=251i=1∑25E(Xi)=2525μ=μ
So Xˉ is unbiased for μ for every n, whatever the population.
The divisor n−1 hits σ2 exactly; the divisor n falls short by the factor nn−1, so it is biased low.
Explain where the missing degree of freedom goes
i=1∑25(Xi−Xˉ)=0
The deviations are forced to sum to zero, so only n−1 of them are free.
Warn that unbiasedness does not survive a square root
E(S2)=σ2butE(S)<σ
The square root is concave, so the square root of an unbiased variance estimator is biased low for σ.
Note that unbiasedness says nothing about variance
E(X1)=μ too, yet Var(X1)=σ2>25σ2=Var(Xˉ)
The first observation alone is also unbiased for μ; Xˉ is preferred because it is far more precise, not because it is unbiased.
Recall what it means for an estimator to be unbiased
θ^ is unbiased for θ⟺E(θ^)=θ
Unbiasedness is a statement about the mean of the estimator's sampling distribution, not about any one sample.
Show that the sample mean is unbiased for μ
E(Xˉ)=n1i=1∑nE(Xi)=nnμ=μ
Expectation is linear, so the n copies of μ divide out exactly.
Recall the unbiased estimator of the population variance
S2=n−11i=1∑n(Xi−Xˉ)2,E(S2)=σ2
The divisor n−1 is what makes the expectation come out as σ2 exactly.
Recall why the divisor n gives a biased estimator
E(n1i=1∑n(Xi−Xˉ)2)=nn−1σ2<σ2
The deviations are taken about Xˉ rather than μ, so they are too small on average; dividing by n leaves the estimate biased low.
Recall the computational form of the sum of squares
Sxx=∑x2−n(∑x)2
This is the identity that turns the summary statistics into the sum of squared deviations without listing the data.
Recall the variance of the sample mean
Var(Xˉ)=nσ2
The observations are independent, so the variances add and the factor n1 is squared.
Recall the standard error of the mean
se(Xˉ)=nσ
It is σ divided by n, not by n.
Recall the degrees of freedom of the t-distribution used here
ν=n−1
One degree of freedom is spent estimating μ by xˉ.
Recall why t and not z is used when σ is unknown
S/nXˉ−μ∼tn−1,tν>z for every finite ν
The t-distribution has heavier tails, so it gives a wider (more honest) interval.
Select the statement that follows from these expectations
E(Xˉ)=μ,E(S2)=σ2,\qquadE(25Sxx)=2524σ2
These three expectations settle which estimators are unbiased and which is not.
Answer
E(Xˉ)=μ,E(S2)=σ2
Question 2
9 markschallenging
A confidence interval for the mean reaction time μ of trial responses is calculated from a random sample of size n, using a known population standard deviation σ. Which of the following statements about the width of a confidence interval is correct?
Show worked solution
Worked solution
Write down the width of the interval
width=2znσ
Everything about the width is contained in this one expression.
Read off the dependence on the sample size
width∝n1
Only n appears, so quadrupling n halves the width.
Read off the dependence on the confidence level
z90%=1.6449<z95%=1.9600<z99%=2.5758
A higher confidence level requires a bigger percentage point, hence a wider interval.
Test the effect of quadrupling the sample size
2zσ/n2zσ/4n=21
Four times the data gives half the width, not a quarter of it.
Test the effect of doubling the sample size
2zσ/n2zσ/2n=21≈0.7071
Doubling n multiplies the width by about 0.71; it does not halve it.
Note that the width does not depend on the data
width does not involve xˉ
With σ known, the width is fixed before the sample is even taken.
Recall what it means for an estimator to be unbiased
θ^ is unbiased for θ⟺E(θ^)=θ
Unbiasedness is a statement about the mean of the estimator's sampling distribution, not about any one sample.
Show that the sample mean is unbiased for μ
E(Xˉ)=n1i=1∑nE(Xi)=nnμ=μ
Expectation is linear, so the n copies of μ divide out exactly.
Recall the unbiased estimator of the population variance
S2=n−11i=1∑n(Xi−Xˉ)2,E(S2)=σ2
The divisor n−1 is what makes the expectation come out as σ2 exactly.
Recall why the divisor n gives a biased estimator
E(n1i=1∑n(Xi−Xˉ)2)=nn−1σ2<σ2
The deviations are taken about Xˉ rather than μ, so they are too small on average; dividing by n leaves the estimate biased low.
Recall the computational form of the sum of squares
Sxx=∑x2−n(∑x)2
This is the identity that turns the summary statistics into the sum of squared deviations without listing the data.
Recall the variance of the sample mean
Var(Xˉ)=nσ2
The observations are independent, so the variances add and the factor n1 is squared.
Recall the standard error of the mean
se(Xˉ)=nσ
It is σ divided by n, not by n.
Recall the degrees of freedom of the t-distribution used here
ν=n−1
One degree of freedom is spent estimating μ by xˉ.
Select the statement consistent with the width formula
width=2znσ⇒width↑ with z,width↓ with n
The width grows with the confidence level and shrinks like n1.
Answer
width=2znσ
Question 3
9 markschallenging
A random sample of batteries gives a 98% confidence interval for the mean lifetime μ, in hours, of (215.4,224.6). Which of the following is the correct interpretation of this confidence interval?
Show worked solution
Worked solution
Identify what is random and what is fixed
μ is a fixed constant;(lower,upper) is random
The endpoints depend on the sample, so they change from sample to sample; μ does not.
State the property the construction actually guarantees
P(Xˉ−cse<μ<Xˉ+cse)=0.98
The probability statement is made BEFORE the sample is taken, about the random interval.
Explain why the same statement cannot be made afterwards
P(215.4<μ<224.6)∈{0,1}
Once the numbers are computed, μ either is or is not inside; no randomness is left to carry a probability.
Rule out the interpretation about individual observations
CI for μ=interval containing 98% of the data
A confidence interval is about the population MEAN, not about where individual values fall.
Rule out the interpretation about the sample mean
xˉ=2215.4+224.6∈(215.4,224.6) always
xˉ is the midpoint, so it is inside with probability 1; saying so tells us nothing.
Express the guarantee as a long-run frequency
proportion of intervals containing μ⟶0.98
Repeating the sampling many times, this is the fraction of the resulting intervals that capture μ.
Recall what it means for an estimator to be unbiased
θ^ is unbiased for θ⟺E(θ^)=θ
Unbiasedness is a statement about the mean of the estimator's sampling distribution, not about any one sample.
Show that the sample mean is unbiased for μ
E(Xˉ)=n1i=1∑nE(Xi)=nnμ=μ
Expectation is linear, so the n copies of μ divide out exactly.
Recall the unbiased estimator of the population variance
S2=n−11i=1∑n(Xi−Xˉ)2,E(S2)=σ2
The divisor n−1 is what makes the expectation come out as σ2 exactly.
Recall why the divisor n gives a biased estimator
E(n1i=1∑n(Xi−Xˉ)2)=nn−1σ2<σ2
The deviations are taken about Xˉ rather than μ, so they are too small on average; dividing by n leaves the estimate biased low.
Recall the computational form of the sum of squares
Sxx=∑x2−n(∑x)2
This is the identity that turns the summary statistics into the sum of squared deviations without listing the data.
Recall the variance of the sample mean
Var(Xˉ)=nσ2
The observations are independent, so the variances add and the factor n1 is squared.
Recall the standard error of the mean
se(Xˉ)=nσ
It is σ divided by n, not by n.
Recall the degrees of freedom of the t-distribution used here
ν=n−1
One degree of freedom is spent estimating μ by xˉ.
Recall why t and not z is used when σ is unknown
S/nXˉ−μ∼tn−1,tν>z for every finite ν
The t-distribution has heavier tails, so it gives a wider (more honest) interval.
Recall the limiting behaviour of the t-distribution
tν→N(0,1)as ν→∞
For a large sample the t value and the z value are almost the same, but for a small sample they are not.
Select the statement describing the long-run behaviour of the method
98% of such intervals contain μ
This is the only one of the five statements that is true of a confidence interval.
Answer
98% of intervals constructed in this way contain μ
Question 4
9 markschallenging
The volume, in millilitres, of a bottle of cordial from the old process is normally distributed with mean μ1, and from the new process it is normally distributed with mean μ2. The two population variances are unknown but may be assumed equal. Independent random samples give n1=14, xˉ1=37.6, s12=5.5 and n2=16, xˉ2=33.9, s22=6.2. The percentage point of the t-distribution with ν=28 degrees of freedom required here is t=2.048. Which of the following is the 95% confidence interval for μ1−μ2, with each endpoint given to 4 decimal places?
Show worked solution
Worked solution
Find the point estimate of μ1−μ2
xˉ1−xˉ2=37.6−33.9=3.7
The difference of the two unbiased sample means is unbiased for μ1−μ2.
Both variances are unknown, so they must be pooled and a t percentage point used.
Find the degrees of freedom
ν=n1+n2−2=14+16−2=28
Each sample loses one degree of freedom to its own mean.
Check that the pooled estimate lies between the two estimates
5.5≤5.875≤6.2
A weighted mean of the two variance estimates must lie between them.
Form the lower endpoint
3.7000−1.8166=1.8834
Subtract the margin of error from the point estimate of the difference.
Form the upper endpoint
3.7000+1.8166=5.5166
Add the margin of error to the point estimate of the difference.
Check the midpoint of the interval
21.8834+5.5166=3.7000
The interval is symmetric about xˉ1−xˉ2.
Read off what the interval says about the two means
0∈/(1.8834,5.5166)
Zero does not lie in the interval, so at this confidence level the data are not consistent with equal population means.
Check the width of the interval
5.5166−1.8834=3.6333
The width is exactly twice the margin of error.
Recall what it means for an estimator to be unbiased
θ^ is unbiased for θ⟺E(θ^)=θ
Unbiasedness is a statement about the mean of the estimator's sampling distribution, not about any one sample.
State the confidence interval for the difference
3.7000±1.8166=(1.8834,5.5166)
This is the required 95% confidence interval for μ1−μ2.
Answer
95% CI=(1.8834,5.5166)
Question 5
9 markschallenging
The length, in millimetres, of a steel pin is normally distributed with unknown mean μ and known standard deviation σ=4.5. A 95% confidence interval for μ is to be found from a random sample of size n. The percentage point of the standard normal distribution required here is z=1.9600. Find the smallest sample size n for which the width of the confidence interval is at most w=2.5.
Show worked solution
Worked solution
Write down the width of a known-variance confidence interval
width=2znσ
The interval reaches zσ/n either side of xˉ.
Write down the requirement as an inequality
2×1.9600×n4.5≤2.5
The width must not exceed the value stated in the question.
Rearrange to isolate n
n≥w2zσ=2.52×1.9600×4.5=7.0560
Every quantity here is positive, so the inequality does not reverse.
Square both sides
n≥(w2zσ)2=49.7871
Squaring is valid because both sides are positive.
Round up to the next whole number
n=⌈49.7871⌉=50
Rounding DOWN would leave the interval too wide, so the ceiling is required.
Check that this sample size meets the requirement
2×1.9600×504.5=2.4947≤2.5
The width achieved with n=50 is within the limit.
Check that one fewer observation would not
2×1.9600×494.5=2.5200>2.5
So n=50 really is the smallest sample size that works.
Note that the requirement does not involve the data
n depends only on z,σ and w
With σ known, the width is settled before a single observation is taken.
Recall what it means for an estimator to be unbiased
θ^ is unbiased for θ⟺E(θ^)=θ
Unbiasedness is a statement about the mean of the estimator's sampling distribution, not about any one sample.
Show that the sample mean is unbiased for μ
E(Xˉ)=n1i=1∑nE(Xi)=nnμ=μ
Expectation is linear, so the n copies of μ divide out exactly.
Recall the unbiased estimator of the population variance
S2=n−11i=1∑n(Xi−Xˉ)2,E(S2)=σ2
The divisor n−1 is what makes the expectation come out as σ2 exactly.
Recall why the divisor n gives a biased estimator
E(n1i=1∑n(Xi−Xˉ)2)=nn−1σ2<σ2
The deviations are taken about Xˉ rather than μ, so they are too small on average; dividing by n leaves the estimate biased low.
Recall the computational form of the sum of squares
Sxx=∑x2−n(∑x)2
This is the identity that turns the summary statistics into the sum of squared deviations without listing the data.
Recall the variance of the sample mean
Var(Xˉ)=nσ2
The observations are independent, so the variances add and the factor n1 is squared.
State the smallest sample size
n=50
This is the least number of observations achieving the required width.
Answer
n=50
Unlock 29 more Estimation and confidence intervals questions
Create a free account to work through every Further Maths Estimation and confidence intervals question with instant step-by-step worked solutions, progress tracking and interactive lessons.