Skip to the content

Topics Covered

Meaning Types Scatter Diagram Karl Pearson's r Rank Correlation Tied Ranks Properties Bivariate Frequency
On this page
  1. 1. Meaning of Correlation
  2. 2. Types of Correlation
  3. 3. Scatter Diagram
  4. 4. Karl Pearson's Coefficient of Correlation
  5. 5. Spearman's Rank Correlation Coefficient
  6. 6. Correlation for Bivariate Frequency Distribution
  7. Key Take-aways

1. Meaning of Correlation

DEFINITION

Two variables are said to be correlated if a change in one is accompanied by a systematic change in the other. Correlation measures the degree and direction of association between two variables.

Croxton & Cowden: "When the relationship is of a quantitative nature, the appropriate statistical tool for discovering and measuring the relationship is known as correlation."

2. Types of Correlation

3. Scatter Diagram

DEFINITION

A scatter diagram plots the bivariate observations \((x_i, y_i)\) as points in the Cartesian plane. The pattern of points reveals the direction, strength and form of the relationship.

Strong Positive (r ≈ +0.95)
r close to +1
Strong Negative (r ≈ −0.95)
r close to −1
No Correlation (r ≈ 0)
random scatter

4. Karl Pearson's Coefficient of Correlation

DEFINITION

For \(n\) bivariate observations, the Karl Pearson correlation coefficient is

\[ r \;=\; \dfrac{\text{Cov}(X, Y)}{\sigma_X \sigma_Y} \;=\; \dfrac{\sum (x_i - \bar x)(y_i - \bar y)}{\sqrt{\sum (x_i - \bar x)^2 \cdot \sum (y_i - \bar y)^2}}. \]

Computational form

\[ r \;=\; \dfrac{n \sum xy - \sum x \sum y}{\sqrt{[n \sum x^2 - (\sum x)^2][n \sum y^2 - (\sum y)^2]}}. \]

Properties of \(r\)

  1. \(-1 \le r \le 1\).
  2. \(r = +1\): perfect positive linear correlation.
  3. \(r = -1\): perfect negative linear correlation.
  4. \(r = 0\): no linear correlation (variables may still be related non-linearly).
  5. Independent of change of origin and scale: if \(u = (x - a)/h,\; v = (y - b)/k\) then \(r_{xy} = r_{uv}\) (assuming \(h, k\) have the same sign).
  6. Symmetric: \(r_{xy} = r_{yx}\).
  7. Geometric mean of regression coefficients: \(r = \pm \sqrt{b_{yx}\, b_{xy}}\).

Interpretation

|r|Strength
0.0 – 0.3Weak / negligible
0.3 – 0.7Moderate
0.7 – 1.0Strong

Worked Examples

EXAMPLE 1

Compute \(r\) for the data:

x12345
y25387

\(\sum x = 15,\; \sum y = 25,\; \sum xy = 88,\; \sum x^2 = 55,\; \sum y^2 = 151,\; n = 5\).

Numerator: \(5(88) - 15(25) = 440 - 375 = 65\).

Denominator: \(\sqrt{(5 \cdot 55 - 225)(5 \cdot 151 - 625)} = \sqrt{50 \cdot 130} = \sqrt{6500} = 80.62\).

\(r = 65 / 80.62 = 0.806\) — strong positive correlation.

EXAMPLE 2 (Change of origin/scale)

Heights (cm) of 5 men and weights (kg): use \(u = x - 170,\; v = y - 60\). Sums: \(\sum u = 5,\; \sum v = -3,\; \sum uv = 67,\; \sum u^2 = 110,\; \sum v^2 = 51,\; n = 5\).

\(r_{uv} = (5 \cdot 67 - 5 \cdot (-3))/\sqrt{(5\cdot 110 - 25)(5\cdot 51 - 9)} = (335 + 15)/\sqrt{525 \cdot 246} = 350/359.37 = 0.974\).

By the property, \(r_{xy} = r_{uv} = 0.974\).

5. Spearman's Rank Correlation Coefficient

DEFINITION

When data are ranked (e.g. judges' ratings), Spearman's coefficient is

\[ \rho \;=\; 1 - \dfrac{6 \sum d_i^2}{n(n^2 - 1)}, \]

where \(d_i = R_{1i} - R_{2i}\) is the difference between the two sets of ranks for the \(i\)-th individual.

It is the Pearson correlation applied to ranks. \(-1 \le \rho \le 1\).

5.1 With Ties

If some observations have the same rank, assign each tied observation the average of the ranks they would otherwise have occupied. Then use the corrected formula:

\[ \rho \;=\; 1 - \dfrac{6 \left[\sum d_i^2 + \sum \dfrac{m_j^3 - m_j}{12}\right]}{n(n^2 - 1)}, \]

where \(m_j\) is the number of items in the \(j\)-th tie group, and the sum is over all tie groups.

EXAMPLE 1 (Without ties)

Two judges rank 5 contestants:

ContestantABCDE
Judge 113254
Judge 224153
d−1−1101
d²11101

\(\sum d^2 = 4;\; \rho = 1 - 6(4)/(5(24)) = 1 - 24/120 = 0.80\).

EXAMPLE 2 (With ties)

Marks of 5 students: X = 80, 85, 85, 70, 90; Y = 70, 80, 90, 60, 95.

Ranks of X: 4, 2.5, 2.5, 5, 1 (tie at 85, ranks 2 & 3 averaged → 2.5; tie group size m = 2).

Ranks of Y: 4, 3, 2, 5, 1 (no ties).

\(d\): 0, −0.5, 0.5, 0, 0; \(\sum d^2 = 0.5\). Tie correction: \((2^3 - 2)/12 = 0.5\).

\(\rho = 1 - 6(0.5 + 0.5)/(5 \cdot 24) = 1 - 6/120 = 0.95\).

6. Correlation for Bivariate Frequency Distribution

If data are tabulated in a two-way table (rows for \(X\), columns for \(Y\)) with frequencies \(f_{ij}\):

\[ r \;=\; \dfrac{N \sum f xy - \sum fx \sum fy}{\sqrt{\bigl[N \sum f x^2 - (\sum fx)^2\bigr]\bigl[N \sum f y^2 - (\sum fy)^2\bigr]}}. \]

Where \(N = \sum f\) and \(\sum f xy = \sum_{i,j} f_{ij} x_i y_j\).

Step deviation form

Let \(u = (x - A)/h, v = (y - B)/k\). Then by change of origin/scale property:

\[ r_{xy} \;=\; r_{uv} \;=\; \dfrac{N \sum fuv - \sum fu \sum fv}{\sqrt{\bigl[N \sum f u^2 - (\sum fu)^2\bigr]\bigl[N \sum f v^2 - (\sum fv)^2\bigr]}}. \]
EXAMPLE 1

For a small bivariate table with totals: \(N = 50,\; \sum fu = -10,\; \sum fv = 5,\; \sum fuv = 25,\; \sum fu^2 = 90,\; \sum fv^2 = 70\):

Numerator = \(50(25) - (-10)(5) = 1250 + 50 = 1300\).

Denominator = \(\sqrt{(50 \cdot 90 - 100)(50 \cdot 70 - 25)} = \sqrt{4400 \cdot 3475} = \sqrt{15290000} = 3910.2\).

\(r = 1300/3910.2 = 0.332\) (moderate positive).

EXAMPLE 2

If a marketing analyst finds for advertising-vs-sales bivariate frequency table that step-deviation sums give \(r_{uv} = 0.85\), then since change of origin/scale doesn't affect \(r\), \(r_{xy} = 0.85\) — strong positive correlation.

Key Take-aways

Extra Practical Problems

PRACTICE

Additional worked problems with step-by-step procedures to support self-study, matching this unit's topics.

STEP-BY-STEP PROCEDURE
  1. Compute the means \(\bar X\) and \(\bar Y\).
  2. Form the deviations \((X_i-\bar X)\) and \((Y_i-\bar Y)\); then the products \((X_i-\bar X)(Y_i-\bar Y)\) and the squares \((X_i-\bar X)^2,\ (Y_i-\bar Y)^2\); total each column.
  3. Apply Karl Pearson's coefficient \(r = \dfrac{\sum (X_i-\bar X)(Y_i-\bar Y)}{\sqrt{\sum (X_i-\bar X)^2\ \sum (Y_i-\bar Y)^2}}\) (equivalently the product-moment form with \(\sum XY,\ \sum X,\ \sum Y\)).
  4. Interpret: \(r\in[-1,1]\); sign gives direction, magnitude gives strength. \(r\) is independent of change of origin and scale.
  5. Test of significance (\(H_0: \rho = 0\)): \(t = \dfrac{r\sqrt{n-2}}{\sqrt{1-r^2}}\) with \(n-2\) df. If \(t_{cal} > t_{tab}\), the correlation is significant.

Problem 1 — Karl Pearson's r and its Test of Significance

DATA

Marks of 8 students in Mathematics (X) and Statistics (Y):

X2530323537404245
Y810151720222425

The means are \(\bar X = 286/8 = 35.75\) and \(\bar Y = 141/8 = 17.625\). Taking deviations from these exact means, \(\sum(X-\bar X)(Y-\bar Y) = 287.25\), \(\sum(X-\bar X)^2 = 307.5\), \(\sum(Y-\bar Y)^2 = 277.875\).

A shortcut, and its correction. Deviations from the rounded values 36 and 18 are easier, \(d_x = X - 36\), \(d_y = Y - 18\), and give \(\sum d_xd_y = 288\), \(\sum d_x^2 = 308\), \(\sum d_y^2 = 279\) — but those are sums about an assumed mean, not the actual one, and must be corrected: with \(\sum d_x = -2\), \(\sum d_y = -3\), \(n = 8\), \(288 - (-2)(-3)/8 = 287.25\), \(308 - (-2)^2/8 = 307.5\), \(279 - (-3)^2/8 = 277.875\).

\(r = \dfrac{\sum(X-\bar X)(Y-\bar Y)}{\sqrt{\sum(X-\bar X)^2\,\sum(Y-\bar Y)^2}} = \dfrac{287.25}{\sqrt{307.5\times277.875}} = 0.983.\)

Test of significance: \(t = \dfrac{r\sqrt{n-2}}{\sqrt{1-r^2}} = \dfrac{0.9827\sqrt{6}}{\sqrt{1-0.9827^2}} = 12.99\) at 6 df. Table \(t_{0.05,6} = 2.447\). Since \(12.99 > 2.447\), \(H_0\) is rejected — the correlation is highly significant.

Problem 2 — Corrected Correlation Coefficient

DATA

For 50 pairs: \(\bar X = 10, \sigma_x = 3, \bar Y = 6, \sigma_y = 2, r = 0.3\). One inaccurate pair \((X=10, Y=6)\) is removed, leaving 49 pairs. Find the corrected \(r\).

Method: recompute corrected \(\sum X = 500-10 = 490\) (so corrected \(\bar X = 490/49 = 10\), unchanged); similarly corrected \(\bar Y = 6\). Update \(\sum X^2, \sum Y^2, \sum XY\) by removing the squares and the cross-product of the deleted pair, then re-apply \(r = \dfrac{n\sum XY - \sum X\sum Y}{\sqrt{n\sum X^2 - (\sum X)^2}\,\sqrt{n\sum Y^2 - (\sum Y)^2}}\).

Result. The removed pair sat exactly at both means, so its deviations \(x - \bar X\) and \(y - \bar Y\) are both zero: removing it takes nothing out of \(\sum(x-\bar x)^2\), \(\sum(y-\bar y)^2\) or \(\sum(x-\bar x)(y-\bar y)\). The standard deviations change (the same sums now divide by 49), but \(r\) is a ratio in which that divisor cancels, so the corrected \(r\) is exactly 0.3.