Hard A-Level Correlation and regression Questions

Challenging, exam-style A-Level Correlation and regression questions with worked solutions. Stretch yourself on the hardest regression line, prediction, substitution, inverse prediction problems.

regression linepredictionsubstitutioninverse predictionrearrangegradient
A-Level34 questionsStep-by-step solutions
Question 1
8 markschallenging
A shop tracks heaters in winter. They investigate how the number of heaters sold depends on the outside temperature (°C). Plotting xx on the horizontal axis and yy on the vertical axis, which statement correctly identifies the variables?
Show worked solution

Worked solution

  1. Decide which variable is chosen or controlled.

    x=explanatoryx = \text{explanatory}

    The explanatory (independent) variable is placed on the x-axis.

  2. Decide which variable responds.

    y=responsey = \text{response}

    The response (dependent) variable is placed on the y-axis.

  3. Match the variables to this context.

    contextx,y\text{context} \to x, y

    Assign the roles using the wording of the problem.

  4. Interpret the gradient in context.

    b=2.2b = -2.2

    For each extra unit of the outside temperature (°C), the model changes the number of heaters sold by -2.2 heaters.

  5. Interpret the intercept in context.

    a=12a = 12

    When the outside temperature (°C) is 0 the model predicts the number of heaters sold = 12 heaters.

  6. State the valid range of the explanatory variable.

    6x346 \le x \le 34

    Predictions are most trustworthy inside this range.

  7. Distinguish interpolation from extrapolation.

    inside rangeinterpolation\text{inside range} \to \text{interpolation}

    Interpolation is reliable; extrapolation beyond the data is not.

  8. Identify the response variable.

    y is the response variabley \text{ is the response variable}

    The number of heaters sold responds to changes in the explanatory variable.

  9. Identify the explanatory variable.

    x is the explanatory variablex \text{ is the explanatory variable}

    The outside temperature (°c) is the controlled/chosen variable.

  10. Note that correlation does not imply causation.

    correlationcausation\text{correlation} \ne \text{causation}

    A linear relationship alone does not prove one variable causes the other.

  11. Comment on the correlation coefficient.

    r=0.75r = -0.75

    Values close to +1 or -1 indicate strong linear correlation.

  12. Give the answer to a sensible degree of accuracy.

    round to match the data\text{round to match the data}

    Do not quote more precision than the data supports.

  13. Restate the regression equation used.

    y=122.2xy = 12 - 2.2x

    This is the fitted line of y on x.

  14. Check the sign of the gradient.

    b<0b < 0

    The sign of the gradient agrees with the direction of the correlation.

  15. Reflect on the limitations of a linear model.

    relationship may be non-linear\text{relationship may be non-linear}

    The straight-line model may not hold outside the observed data.

Answer
The outside temperature (°c) is the explanatory variable and the number of heaters sold is the response variable
Question 2
8 markschallenging
A researcher studies phone use. For the daily screen time (hours) as xx and the hours of sleep as yy, the product moment correlation coefficient is r=0.9r = -0.9. Which best describes the linear correlation?
Show worked solution

Worked solution

  1. Read the value of the correlation coefficient.

    r=0.9r = -0.9

    The PMCC r summarises the linear correlation.

  2. Classify the magnitude of r.

    rstrength|r| \to \text{strength}

    The closer |r| is to 1, the stronger the correlation.

  3. Determine the sign of r.

    r<0r < 0

    A positive r means positive correlation, negative means negative.

  4. Interpret the gradient in context.

    b=0.8b = -0.8

    For each extra unit of the daily screen time (hours), the model changes the hours of sleep by -0.8 hours.

  5. Interpret the intercept in context.

    a=25a = 25

    When the daily screen time (hours) is 0 the model predicts the hours of sleep = 25 hours.

  6. State the valid range of the explanatory variable.

    0x120 \le x \le 12

    Predictions are most trustworthy inside this range.

  7. Distinguish interpolation from extrapolation.

    inside rangeinterpolation\text{inside range} \to \text{interpolation}

    Interpolation is reliable; extrapolation beyond the data is not.

  8. Identify the response variable.

    y is the response variabley \text{ is the response variable}

    The hours of sleep responds to changes in the explanatory variable.

  9. Identify the explanatory variable.

    x is the explanatory variablex \text{ is the explanatory variable}

    The daily screen time (hours) is the controlled/chosen variable.

  10. Note that correlation does not imply causation.

    correlationcausation\text{correlation} \ne \text{causation}

    A linear relationship alone does not prove one variable causes the other.

  11. Comment on the correlation coefficient.

    r=0.9r = -0.9

    Values close to +1 or -1 indicate strong linear correlation.

  12. Give the answer to a sensible degree of accuracy.

    round to match the data\text{round to match the data}

    Do not quote more precision than the data supports.

  13. Restate the regression equation used.

    y=250.8xy = 25 - 0.8x

    This is the fitted line of y on x.

  14. Check the sign of the gradient.

    b<0b < 0

    The sign of the gradient agrees with the direction of the correlation.

  15. Reflect on the limitations of a linear model.

    relationship may be non-linear\text{relationship may be non-linear}

    The straight-line model may not hold outside the observed data.

Answer
Strong, negative
Question 3
8 markschallenging
An engineer tests engines. The regression line is y=40+3.5xy = 40 + 3.5x, with yy representing the fuel used per 100 km (litres) and xx representing the engine size (litres). What does the intercept represent in this context?
Show worked solution

Worked solution

  1. Identify the intercept a of the line.

    a=40a = 40

    The intercept is the constant term.

  2. Interpret the intercept as the value at x = 0.

    x=0y=ax = 0 \Rightarrow y = a

    It is the predicted y when x is zero.

  3. Express the intercept in the context.

    predicted y when x=0\text{predicted } y \text{ when } x = 0

    State its meaning using the real-world variables.

  4. Interpret the gradient in context.

    b=3.5b = 3.5

    For each extra unit of the engine size (litres), the model changes the fuel used per 100 km (litres) by 3.5 litres.

  5. Interpret the intercept in context.

    a=40a = 40

    When the engine size (litres) is 0 the model predicts the fuel used per 100 km (litres) = 40 litres.

  6. State the valid range of the explanatory variable.

    4x394 \le x \le 39

    Predictions are most trustworthy inside this range.

  7. Distinguish interpolation from extrapolation.

    inside rangeinterpolation\text{inside range} \to \text{interpolation}

    Interpolation is reliable; extrapolation beyond the data is not.

  8. Identify the response variable.

    y is the response variabley \text{ is the response variable}

    The fuel used per 100 km (litres) responds to changes in the explanatory variable.

  9. Identify the explanatory variable.

    x is the explanatory variablex \text{ is the explanatory variable}

    The engine size (litres) is the controlled/chosen variable.

  10. Note that correlation does not imply causation.

    correlationcausation\text{correlation} \ne \text{causation}

    A linear relationship alone does not prove one variable causes the other.

  11. Comment on the correlation coefficient.

    r=0.66r = 0.66

    Values close to +1 or -1 indicate strong linear correlation.

  12. Give the answer to a sensible degree of accuracy.

    round to match the data\text{round to match the data}

    Do not quote more precision than the data supports.

  13. Restate the regression equation used.

    y=40+3.5xy = 40 + 3.5x

    This is the fitted line of y on x.

  14. Check the sign of the gradient.

    b>0b > 0

    The sign of the gradient agrees with the direction of the correlation.

  15. Reflect on the limitations of a linear model.

    relationship may be non-linear\text{relationship may be non-linear}

    The straight-line model may not hold outside the observed data.

Answer
The predicted the fuel used per 100 km (litres) when the engine size (litres) is 0, namely 40 litres
Question 4
8 markschallenging
A taxi firm logs journeys. The regression line of yy on xx is y=15+1.2xy = 15 + 1.2x, where yy is the fare charged (£) and xx is the journey distance (km). What does the gradient represent in this context?
Show worked solution

Worked solution

  1. Identify the gradient b of the line.

    b=1.2b = 1.2

    The gradient is the coefficient of x.

  2. Interpret the gradient as a rate of change.

    b=ΔyΔxb = \frac{\Delta y}{\Delta x}

    It measures how much y changes per unit change in x.

  3. Express that rate in the context.

    change in y per unit x\text{change in } y \text{ per unit } x

    State the meaning using the real-world variables.

  4. Interpret the gradient in context.

    b=1.2b = 1.2

    For each extra unit of the journey distance (km), the model changes the fare charged (£) by 1.2 pounds.

  5. Interpret the intercept in context.

    a=15a = 15

    When the journey distance (km) is 0 the model predicts the fare charged (£) = 15 pounds.

  6. State the valid range of the explanatory variable.

    3x253 \le x \le 25

    Predictions are most trustworthy inside this range.

  7. Distinguish interpolation from extrapolation.

    inside rangeinterpolation\text{inside range} \to \text{interpolation}

    Interpolation is reliable; extrapolation beyond the data is not.

  8. Identify the response variable.

    y is the response variabley \text{ is the response variable}

    The fare charged (£) responds to changes in the explanatory variable.

  9. Identify the explanatory variable.

    x is the explanatory variablex \text{ is the explanatory variable}

    The journey distance (km) is the controlled/chosen variable.

  10. Note that correlation does not imply causation.

    correlationcausation\text{correlation} \ne \text{causation}

    A linear relationship alone does not prove one variable causes the other.

  11. Comment on the correlation coefficient.

    r=0.88r = 0.88

    Values close to +1 or -1 indicate strong linear correlation.

  12. Give the answer to a sensible degree of accuracy.

    round to match the data\text{round to match the data}

    Do not quote more precision than the data supports.

  13. Restate the regression equation used.

    y=15+1.2xy = 15 + 1.2x

    This is the fitted line of y on x.

  14. Check the sign of the gradient.

    b>0b > 0

    The sign of the gradient agrees with the direction of the correlation.

  15. Reflect on the limitations of a linear model.

    relationship may be non-linear\text{relationship may be non-linear}

    The straight-line model may not hold outside the observed data.

Answer
On average, the fare charged (£) changes by 1.2 pounds for each 1-unit increase in the journey distance (km)
Question 5
8 markschallenging
A gym records member data. A strong correlation is found between the weekly training hours and the distance run in a test (km). Can we conclude that changes in xx cause the changes in yy?
Show worked solution

Worked solution

  1. Recall what correlation measures.

    correlation: linear association\text{correlation: linear association}

    Correlation only measures how points cluster about a line.

  2. Recall the causation principle.

    associationcause\text{association} \ne \text{cause}

    Association does not establish a causal mechanism.

  3. Apply the principle to this context.

    possible lurking variable\text{possible lurking variable}

    A third factor could drive both variables.

  4. Interpret the gradient in context.

    b=5b = 5

    For each extra unit of the weekly training hours, the model changes the distance run in a test (km) by 5 km.

  5. Interpret the intercept in context.

    a=8a = 8

    When the weekly training hours is 0 the model predicts the distance run in a test (km) = 8 km.

  6. State the valid range of the explanatory variable.

    0x180 \le x \le 18

    Predictions are most trustworthy inside this range.

  7. Distinguish interpolation from extrapolation.

    inside rangeinterpolation\text{inside range} \to \text{interpolation}

    Interpolation is reliable; extrapolation beyond the data is not.

  8. Identify the response variable.

    y is the response variabley \text{ is the response variable}

    The distance run in a test (km) responds to changes in the explanatory variable.

  9. Identify the explanatory variable.

    x is the explanatory variablex \text{ is the explanatory variable}

    The weekly training hours is the controlled/chosen variable.

  10. Note that correlation does not imply causation.

    correlationcausation\text{correlation} \ne \text{causation}

    A linear relationship alone does not prove one variable causes the other.

  11. Comment on the correlation coefficient.

    r=0.7r = 0.7

    Values close to +1 or -1 indicate strong linear correlation.

  12. Give the answer to a sensible degree of accuracy.

    round to match the data\text{round to match the data}

    Do not quote more precision than the data supports.

  13. Restate the regression equation used.

    y=8+5xy = 8 + 5x

    This is the fitted line of y on x.

  14. Check the sign of the gradient.

    b>0b > 0

    The sign of the gradient agrees with the direction of the correlation.

  15. Reflect on the limitations of a linear model.

    relationship may be non-linear\text{relationship may be non-linear}

    The straight-line model may not hold outside the observed data.

Answer
No - a strong correlation does not by itself show that one variable causes the other

Unlock 29 more Correlation and regression questions

Create a free account to work through every A-Level Correlation and regression question with instant step-by-step worked solutions, progress tracking and interactive lessons.

  • Full worked solutions for every question
  • Interactive lessons and instant feedback
  • Track your mastery across every topic
Create a Free Account

No card required · Free forever

More Correlation and regression practice

Related Statistics topics