Correlation & Linear Regression Calculator

Correlation measures how strongly two variables move together, and regression finds the straight line that best fits them. Enter paired x and y values to get r, R², the regression equation and a prediction.

How it is calculated

Enter x values and matching y values.

Optionally enter an x to predict.

Read r, R² and the regression line.

Formula

Slope: b = Σ(x − x̄)(y − ȳ) ÷ Σ(x − x̄)²

Intercept: a = ȳ − b x̄

Correlation: r = Sxy ÷ √(Sxx × Syy)

What is the Correlation & Linear Regression Calculator?

This calculator studies paired data, two measurements taken on the same things, such as hours studied and marks scored for each student. It gives Pearson's correlation coefficient r, which measures how closely the points follow a straight line; R², the share of variation in y that the line explains; the least-squares regression line y = a + bx; and a predicted y for any x you choose.

Correlation and regression are among the most used tools in science, economics and business. They show up in Class 11 and Class 12 statistics, economics projects, lab work where you fit a calibration line, and analytics jobs where you predict sales from advertising spend. The calculator shows the intermediate sums Sxy, Sxx and Syy, which is how exam solutions are set out.

How to calculate it by hand

1. Find the means x̄ and ȳ.

2. For each pair, compute the deviations x − x̄ and y − ȳ.

3. Compute Sxy = Σ(x − x̄)(y − ȳ), Sxx = Σ(x − x̄)² and Syy = Σ(y − ȳ)².

4. Slope: b = Sxy ÷ Sxx. Intercept: a = ȳ − b x̄.

5. Correlation: r = Sxy ÷ √(Sxx × Syy). Square it for R².

6. Predict: substitute the new x into y = a + bx.

Least squares: the line with the smallest total error

For any candidate line y = a + bx, each data point has a vertical error, its actual y minus the line's prediction. Least squares chooses a and b to make the sum of the squared errors as small as possible. Setting the partial derivatives with respect to a and b to zero gives two normal equations whose solution is b = Sxy ÷ Sxx and a = ȳ − b x̄. The second result shows the fitted line always passes through the point (x̄, ȳ). Squaring the errors penalises large misses heavily, which is why one extreme point can tilt the line.

What r measures and why it lies between −1 and 1

Sxy is positive when points above the mean in x also tend to be above the mean in y, and negative when they are opposite. Dividing by √(Sxx × Syy) removes the units and scales the result, and the Cauchy–Schwarz inequality guarantees it stays between −1 and +1. r = ±1 means every point lies exactly on a line. r near 0 means no linear pattern, although a strong curved relationship, like a U-shape, can still give r = 0. R² = r² reports the fraction of the variation in y accounted for by the line.

Two regression lines, and correlation is not causation

The regression line of y on x minimises vertical errors and is used to predict y. The line of x on y minimises horizontal errors and is a different line unless r = ±1; the two slopes multiply to r². Use the one that matches what you want to predict. Also, a strong r does not show that x causes y. Ice-cream sales and drowning cases both rise in summer. Finally, predictions far outside the range of the data, called extrapolation, can go badly wrong because the linear pattern may not continue.

Worked example, step by step

Ms Iyer, a maths teacher in Coimbatore, recorded weekly self-study hours and test marks for six students: hours 2, 3, 5, 6, 8, 9 and marks 45, 52, 60, 66, 75, 80. She wants the line of best fit and a predicted mark for a student who studies 7 hours.

Means: x̄ = 5.5, ȳ = 63

Sums of squares: Sxy = Σ(x − x̄)(y − ȳ) = 183 Sxx = 37.5, Syy = 896

Slope and intercept: b = Sxy ÷ Sxx = 4.88 a = ȳ − b x̄ = 36.16

Correlation: r = Sxy ÷ √(Sxx × Syy) = 0.998347, R² = 0.996696

Prediction: y = 36.16 + 4.88 × 7 = 70.32

Answer: Correlation (r) 0.9983; Regression line y = 36.16 + 4.88x; R² 0.9967

Common mistakes to avoid

Entering x and y lists in different orders, which breaks the pairing and gives a meaningless r.

Reading a high r as proof of cause and effect.

Using the y-on-x line to predict x from y.

Predicting far outside the observed range of x.

Ignoring a single outlier that is driving most of the correlation; always plot the points first.

Where it is used

Fitting calibration lines in chemistry and physics labs, such as absorbance against concentration.

Economics projects relating income and consumption, or price and demand.

Business forecasts of sales from advertising spend or footfall.

Checking whether two exam scores or two test methods agree.

Estimating trends in agriculture, such as yield against fertiliser dose within a tested range.

Frequently asked questions

Does a high correlation mean one variable causes the other?

No. Correlation shows association, not causation.

What does r = 0 mean?

There is no linear relationship, though a curved relationship may still exist.