Correlation & least-squares regression
The correlation coefficient r measures LINEAR association on a scale from −1 to 1; r² is the proportion of variance in y explained by the model. The least-squares line minimises the sum of squared vertical residuals, has slope b = r·(s_y/s_x), and always passes through (x̄, ȳ) — a fact worth using as a check.
✓ Unlimited questions · marked criterion by criterion · no card needed
Method: how to approach it
The order below is what examiners expect to see, and each step carries its own marks.
- Plot firstA scatterplot reveals curvature, outliers and clusters that r cannot detect. Anscombe’s quartet exists to make this point.
- Compute r and interpret its sign and strengthNear ±1 is strong linear association; near 0 means no LINEAR relation, which is not the same as no relation.
- Fit the lineb = r·s_y/s_x, then a = ȳ − b·x̄ so the line passes through the mean point.
- Interpret the slope in context and check residualsThe slope is the predicted change in y per unit increase in x. Patterned residuals mean the linear model is wrong.
Worked example
With x̄ = 10, ȳ = 50, s_x = 2, s_y = 6 and r = 0.8, find the regression line.
- Slope: b = r·s_y/s_x = 0.8 × 6/2.
- = 2.4.
- Intercept: a = ȳ − b·x̄ = 50 − 2.4 × 10 = 26.
- Check: at x = 10 the line gives 26 + 24 = 50 = ȳ ✓.
Answer. ŷ = 26 + 2.4x, and r² = 0.64 — 64% of the variance in y is explained by x.
Where marks get dropped
These are the specific errors that cost credit on correlation & least-squares regression questions — QED's rubric penalises each of them separately.
- Inferring causation from correlation. A confounder can produce a strong r with no causal link whatever.
- Extrapolating beyond the data range, where the fitted relationship has no support and can be wildly wrong.
- Reporting r ≈ 0 as "no relationship". A perfect parabola has r = 0 while being entirely determined — plot the data.
Practise this until it is automatic
Unlimited fresh questions
QED generates new correlation & least-squares regression problems on demand at warm-up, exam and challenge level, so you can drill this one skill until it stops costing you marks.
Marked like an examiner
Every answer is scored against a point-by-point rubric with partial credit, so you see exactly which step of the method broke down — not just a tick or a cross.
Answer in real notation
A one-tap symbol palette, a visual equation editor and a truth-table builder — or photograph your handwritten working and QED converts it to LaTeX.
Saved to your library
Every question you generate is kept and re-takeable as a timed exam, and your Statistics mastery is tracked so you know when this is exam-ready.
Correlation & least-squares regression — frequently asked questions
What does r² mean?
The fraction of the variance in y explained by the regression on x. r = 0.8 gives r² = 0.64, so 36% of the variation remains unexplained.
Why "least squares"?
The line minimises the sum of squared VERTICAL residuals. Squaring penalises large errors more and yields a unique closed-form solution.
Does the regression line always pass through the mean point?
Yes, (x̄, ȳ) is always on the least-squares line — the fastest check on an intercept calculation.
The rest of Statistics
Describing data, distributions, estimation and hypothesis tests. Each subtopic below has its own method, worked example and mark-losing traps.
- 1Mean, median & mode
- 2Variance, standard deviation & spread
- 3Shape, skew & outliers
- 4Boxplots, histograms & quartiles
- 5The normal distribution & z-scores
- 6Sampling, bias & the sampling distribution
- 7The central limit theorem
- 8Confidence intervals for a mean
- 9Hypothesis testing & p-values
- 10t-tests & comparing two means
- 11Chi-square tests for independence
- 12Correlation & least-squares regression
- 13Type I / Type II errors & power
Ready to make correlation & least-squares regression exam-proof?
Generate your first questions free — no card, no setup, no personal data stored. Practise until the method is second nature.
Start practising free →