Two variables are said to be correlated if a change in one is accompanied by a systematic change in the other. Correlation measures the degree and direction of association between two variables.
Croxton & Cowden: "When the relationship is of a quantitative nature, the appropriate statistical tool for discovering and measuring the relationship is known as correlation."
A scatter diagram plots the bivariate observations \((x_i, y_i)\) as points in the Cartesian plane. The pattern of points reveals the direction, strength and form of the relationship.
For \(n\) bivariate observations, the Karl Pearson correlation coefficient is
\[ r \;=\; \dfrac{\text{Cov}(X, Y)}{\sigma_X \sigma_Y} \;=\; \dfrac{\sum (x_i - \bar x)(y_i - \bar y)}{\sqrt{\sum (x_i - \bar x)^2 \cdot \sum (y_i - \bar y)^2}}. \]| |r| | Strength |
|---|---|
| 0.0 – 0.3 | Weak / negligible |
| 0.3 – 0.7 | Moderate |
| 0.7 – 1.0 | Strong |
Compute \(r\) for the data:
| x | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| y | 2 | 5 | 3 | 8 | 7 |
\(\sum x = 15,\; \sum y = 25,\; \sum xy = 88,\; \sum x^2 = 55,\; \sum y^2 = 151,\; n = 5\).
Numerator: \(5(88) - 15(25) = 440 - 375 = 65\).
Denominator: \(\sqrt{(5 \cdot 55 - 225)(5 \cdot 151 - 625)} = \sqrt{50 \cdot 130} = \sqrt{6500} = 80.62\).
\(r = 65 / 80.62 = 0.806\) — strong positive correlation.
Heights (cm) of 5 men and weights (kg): use \(u = x - 170,\; v = y - 60\). Sums: \(\sum u = 5,\; \sum v = -3,\; \sum uv = 67,\; \sum u^2 = 110,\; \sum v^2 = 51,\; n = 5\).
\(r_{uv} = (5 \cdot 67 - 5 \cdot (-3))/\sqrt{(5\cdot 110 - 25)(5\cdot 51 - 9)} = (335 + 15)/\sqrt{525 \cdot 246} = 350/359.37 = 0.974\).
By the property, \(r_{xy} = r_{uv} = 0.974\).
When data are ranked (e.g. judges' ratings), Spearman's coefficient is
\[ \rho \;=\; 1 - \dfrac{6 \sum d_i^2}{n(n^2 - 1)}, \]where \(d_i = R_{1i} - R_{2i}\) is the difference between the two sets of ranks for the \(i\)-th individual.
It is the Pearson correlation applied to ranks. \(-1 \le \rho \le 1\).
If some observations have the same rank, assign each tied observation the average of the ranks they would otherwise have occupied. Then use the corrected formula:
where \(m_j\) is the number of items in the \(j\)-th tie group, and the sum is over all tie groups.
Two judges rank 5 contestants:
| Contestant | A | B | C | D | E |
|---|---|---|---|---|---|
| Judge 1 | 1 | 3 | 2 | 5 | 4 |
| Judge 2 | 2 | 4 | 1 | 5 | 3 |
| d | −1 | −1 | 1 | 0 | 1 |
| d² | 1 | 1 | 1 | 0 | 1 |
\(\sum d^2 = 4;\; \rho = 1 - 6(4)/(5(24)) = 1 - 24/120 = 0.80\).
Marks of 5 students: X = 80, 85, 85, 70, 90; Y = 70, 80, 90, 60, 95.
Ranks of X: 4, 2.5, 2.5, 5, 1 (tie at 85, ranks 2 & 3 averaged → 2.5; tie group size m = 2).
Ranks of Y: 4, 3, 2, 5, 1 (no ties).
\(d\): 0, −0.5, 0.5, 0, 0; \(\sum d^2 = 0.5\). Tie correction: \((2^3 - 2)/12 = 0.5\).
\(\rho = 1 - 6(0.5 + 0.5)/(5 \cdot 24) = 1 - 6/120 = 0.95\).
If data are tabulated in a two-way table (rows for \(X\), columns for \(Y\)) with frequencies \(f_{ij}\):
Where \(N = \sum f\) and \(\sum f xy = \sum_{i,j} f_{ij} x_i y_j\).
Let \(u = (x - A)/h, v = (y - B)/k\). Then by change of origin/scale property:
For a small bivariate table with totals: \(N = 50,\; \sum fu = -10,\; \sum fv = 5,\; \sum fuv = 25,\; \sum fu^2 = 90,\; \sum fv^2 = 70\):
Numerator = \(50(25) - (-10)(5) = 1250 + 50 = 1300\).
Denominator = \(\sqrt{(50 \cdot 90 - 100)(50 \cdot 70 - 25)} = \sqrt{4400 \cdot 3475} = \sqrt{15290000} = 3910.2\).
\(r = 1300/3910.2 = 0.332\) (moderate positive).
If a marketing analyst finds for advertising-vs-sales bivariate frequency table that step-deviation sums give \(r_{uv} = 0.85\), then since change of origin/scale doesn't affect \(r\), \(r_{xy} = 0.85\) — strong positive correlation.
Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.
Marks of 8 students in Mathematics (X) and Statistics (Y):
| X | 25 | 30 | 32 | 35 | 37 | 40 | 42 | 45 |
|---|---|---|---|---|---|---|---|---|
| Y | 8 | 10 | 15 | 17 | 20 | 22 | 24 | 25 |
The means are \(\bar X = 286/8 = 35.75\) and \(\bar Y = 141/8 = 17.625\). Taking deviations from these exact means, \(\sum(X-\bar X)(Y-\bar Y) = 287.25\), \(\sum(X-\bar X)^2 = 307.5\), \(\sum(Y-\bar Y)^2 = 277.875\).
A shortcut, and its correction. Deviations from the rounded values 36 and 18 are easier, \(d_x = X - 36\), \(d_y = Y - 18\), and give \(\sum d_xd_y = 288\), \(\sum d_x^2 = 308\), \(\sum d_y^2 = 279\) — but those are sums about an assumed mean, not the actual one, and must be corrected: with \(\sum d_x = -2\), \(\sum d_y = -3\), \(n = 8\), \(288 - (-2)(-3)/8 = 287.25\), \(308 - (-2)^2/8 = 307.5\), \(279 - (-3)^2/8 = 277.875\).
\(r = \dfrac{\sum(X-\bar X)(Y-\bar Y)}{\sqrt{\sum(X-\bar X)^2\,\sum(Y-\bar Y)^2}} = \dfrac{287.25}{\sqrt{307.5\times277.875}} = 0.983.\)
Test of significance: \(t = \dfrac{r\sqrt{n-2}}{\sqrt{1-r^2}} = \dfrac{0.9827\sqrt{6}}{\sqrt{1-0.9827^2}} = 12.99\) at 6 df. Table \(t_{0.05,6} = 2.447\). Since \(12.99 > 2.447\), \(H_0\) is rejected — the correlation is highly significant.
For 50 pairs: \(\bar X = 10, \sigma_x = 3, \bar Y = 6, \sigma_y = 2, r = 0.3\). One inaccurate pair \((X=10, Y=6)\) is removed, leaving 49 pairs. Find the corrected \(r\).
Method: recompute corrected \(\sum X = 500-10 = 490\) (so corrected \(\bar X = 490/49 = 10\), unchanged); similarly corrected \(\bar Y = 6\). Update \(\sum X^2, \sum Y^2, \sum XY\) by removing the squares and the cross-product of the deleted pair, then re-apply \(r = \dfrac{n\sum XY - \sum X\sum Y}{\sqrt{n\sum X^2 - (\sum X)^2}\,\sqrt{n\sum Y^2 - (\sum Y)^2}}\).
Result. The removed pair sat exactly at both means, so its deviations \(x - \bar X\) and \(y - \bar Y\) are both zero: removing it takes nothing out of \(\sum(x-\bar x)^2\), \(\sum(y-\bar y)^2\) or \(\sum(x-\bar x)(y-\bar y)\). The standard deviations change (the same sums now divide by 49), but \(r\) is a ratio in which that divisor cancels, so the corrected \(r\) is exactly 0.3.