Skip to the content

Topics Covered

Two-Variable Model Assumptions OLS Estimation Gauss–Markov Partial Correlation Multiple Correlation General Linear Model BLUE
On this page
  1. 1. The Two-Variable Linear Regression Model
  2. 2. Ordinary Least Squares (OLS) Estimation
  3. 3. Gauss–Markov Theorem
  4. 4. Partial and Multiple Correlation
  5. 5. The General Linear (Multiple Regression) Model
  6. 6. Dummy (Indicator) Variables
  7. Key Take-aways from Unit 2

1. The Two-Variable Linear Regression Model

MODEL

The population regression function (PRF):

\[ E(Y \mid X_i) = \beta_0 + \beta_1 X_i. \]

Each individual observation is

\[ Y_i = \beta_0 + \beta_1 X_i + u_i, \]

where \(\beta_0\) is the intercept, \(\beta_1\) is the slope, and \(u_i\) is the stochastic disturbance.

1.1 Classical Assumptions (CLRM)

  1. Linearity in parameters: model linear in \(\beta_0, \beta_1\).
  2. \(X\) is non-stochastic (fixed in repeated sampling), or independent of \(u\).
  3. Zero mean of disturbance: \(E(u_i) = 0\) for all \(i\).
  4. Homoscedasticity: \(\mathrm{Var}(u_i) = \sigma^2\) for all \(i\).
  5. No serial correlation: \(\mathrm{Cov}(u_i, u_j) = 0\) for \(i \ne j\).
  6. \(\mathrm{Cov}(u_i, X_i) = 0\).
  7. \(n > k\) (more observations than parameters).
  8. Variability in \(X\) (not all the same).
  9. Correct functional form.
  10. Normality: \(u_i \sim N(0, \sigma^2)\) — required for hypothesis tests, not for estimation.

2. Ordinary Least Squares (OLS) Estimation

PRINCIPLE

OLS chooses \(\hat\beta_0, \hat\beta_1\) to minimise the residual sum of squares:

\[ \min_{\hat\beta_0,\hat\beta_1}\ \sum_{i=1}^{n} \hat u_i^2 = \sum_{i=1}^{n}(Y_i - \hat\beta_0 - \hat\beta_1 X_i)^2. \]

2.1 The Normal Equations

Setting partial derivatives to zero:

\[ \sum Y_i = n\hat\beta_0 + \hat\beta_1\sum X_i, \qquad \sum X_i Y_i = \hat\beta_0\sum X_i + \hat\beta_1\sum X_i^2. \]

Solving:

\[ \hat\beta_1 = \frac{\sum (X_i - \bar X)(Y_i - \bar Y)}{\sum (X_i - \bar X)^2} = \frac{S_{XY}}{S_{XX}}, \] \[ \hat\beta_0 = \bar Y - \hat\beta_1 \bar X. \]

2.2 Variances of the Estimators

\[ \mathrm{Var}(\hat\beta_1) = \frac{\sigma^2}{S_{XX}}, \qquad \mathrm{Var}(\hat\beta_0) = \sigma^2\left[\frac{1}{n} + \frac{\bar X^2}{S_{XX}}\right], \]

and \(\sigma^2\) is estimated by \(\hat\sigma^2 = \sum \hat u_i^2 /(n-2)\), an unbiased estimator.

EXAMPLE 1 — OLS by hand

Data on 5 households (\(X\) = income, \(Y\) = consumption, ₹'000):

i12345
\(X_i\)1020304050
\(Y_i\)1120273443

\(\bar X = 30, \bar Y = 27\), \(S_{XY} = \sum(X_i-30)(Y_i-27) = 780\), \(S_{XX} = \sum(X_i-30)^2 = 1000\).

\(\hat\beta_1 = 780/1000 = 0.78\), \(\hat\beta_0 = 27 - 0.78\times 30 = 3.60\).

Fitted line: \(\hat Y = 3.60 + 0.78 X\). Marginal propensity to consume is 0.78.

EXAMPLE 2 — Predicting \(Y\)

From Example 1, predict consumption for a household with income ₹35,000: \(\hat Y = 3.60 + 0.78 \times 35 = 30.90\) (i.e. ₹30,900).

For income ₹100,000 (extrapolation beyond data): \(\hat Y = 3.60 + 0.78 \times 100 = 81.60\) (₹81,600). Note: extrapolation is dangerous — the linear relation may not hold outside the sampled range.

3. Gauss–Markov Theorem

STATEMENT

Under the classical assumptions 1–8 (linearity, exogeneity, homoscedasticity, no serial correlation, etc.), the OLS estimators \(\hat\beta_0, \hat\beta_1\) are the Best Linear Unbiased Estimators (BLUE) of the population parameters \(\beta_0, \beta_1\).

3.1 Three Properties Combined: BLUE

Normality of \(u\) is not required for the Gauss–Markov theorem — only for distribution-based inference (\(t\), \(F\) tests).

3.2 Why It Matters

When CLRM assumptions hold, you cannot find a "better" linear unbiased estimator than OLS. Therefore most diagnostics in this course test whether one of the assumptions is violated — and if so, which alternative estimator (GLS, WLS, FGLS, robust SE, etc.) restores BLUE-ness.

4. Partial and Multiple Correlation

4.1 Multiple Correlation \(R\)

For three variables \(Y, X_1, X_2\):

\[ R^2_{Y \cdot X_1 X_2} = \frac{r^2_{YX_1} + r^2_{YX_2} - 2\,r_{YX_1}r_{YX_2}r_{X_1 X_2}}{1 - r^2_{X_1 X_2}}. \]

\(R\) measures the linear association of \(Y\) with the joint influence of \(X_1\) and \(X_2\). It is the correlation between \(Y\) and \(\hat Y\) from the fitted regression.

4.2 Partial Correlation

Partial correlation between \(Y\) and \(X_1\), holding \(X_2\) constant:

\[ r_{YX_1 \cdot X_2} = \frac{r_{YX_1} - r_{YX_2}r_{X_1 X_2}}{\sqrt{(1 - r^2_{YX_2})(1 - r^2_{X_1 X_2})}}. \]

It removes the influence of \(X_2\) on both \(Y\) and \(X_1\) before measuring their association — the "pure" effect of \(X_1\) on \(Y\).

EXAMPLE 1 — Computing partial correlation

\(r_{YX_1} = 0.80, r_{YX_2} = 0.60, r_{X_1 X_2} = 0.50\). Find \(r_{YX_1 \cdot X_2}\).

\(r_{YX_1 \cdot X_2} = (0.80 - 0.60 \times 0.50)/\sqrt{(1-0.36)(1-0.25)} = 0.50/\sqrt{0.48} = 0.50/0.6928 = 0.7217.\)

Net of \(X_2\), \(Y\) and \(X_1\) are still strongly positively associated.

EXAMPLE 2 — Multiple correlation

Using the same numbers:

\(R^2_{Y \cdot X_1 X_2} = (0.64 + 0.36 - 2 \times 0.80 \times 0.60 \times 0.50)/(1 - 0.25) = (1.00 - 0.48)/0.75 = 0.6933.\)

\(R = 0.8326\). Together \(X_1\) and \(X_2\) explain about 69% of the variation in \(Y\).

5. The General Linear (Multiple Regression) Model

With \(k\) explanatory variables:

\[ Y_i = \beta_0 + \beta_1 X_{1i} + \beta_2 X_{2i} + \cdots + \beta_k X_{ki} + u_i, \quad i = 1, \ldots, n. \]

In matrix form:

\[ \mathbf Y = \mathbf X \boldsymbol\beta + \mathbf u, \]

where \(\mathbf Y\) is \(n \times 1\), \(\mathbf X\) is \(n \times (k+1)\), \(\boldsymbol\beta\) is \((k+1) \times 1\), and \(\mathbf u\) is \(n \times 1\).

5.1 OLS in Matrix Form

\[ \hat{\boldsymbol\beta} = (\mathbf X' \mathbf X)^{-1}\,\mathbf X' \mathbf Y. \]

This is the single most important formula in the entire course — every multiple-regression derivation springs from it. The variance-covariance matrix is

\[ \mathrm{Var}(\hat{\boldsymbol\beta}) = \sigma^2 (\mathbf X' \mathbf X)^{-1}. \]

5.2 Properties of the OLS Estimator (BLUE)

5.3 Why "Multiple" Beats "Simple"

EXAMPLE 1 — Wage equation

Mincer wage equation: \(\log W_i = \beta_0 + \beta_1\,\text{educ}_i + \beta_2\,\text{exper}_i + \beta_3\,\text{exper}_i^2 + u_i\). With \(n = 526\) observations from the CPS (US), typical estimates are \(\hat\beta_1 = 0.092\) (each extra year of education raises wages by 9.2%), \(\hat\beta_2 = 0.041\), \(\hat\beta_3 = -0.0007\) (wage rises with experience but at a diminishing rate).

EXAMPLE 2 — Matrix-form computation

For 3 observations \((X, Y) = (1, 2), (2, 4), (3, 5)\), the design matrix (with intercept) is

\(\mathbf X = \begin{pmatrix}1&1\\1&2\\1&3\end{pmatrix}, \mathbf Y = \begin{pmatrix}2\\4\\5\end{pmatrix}\).

\(\mathbf X'\mathbf X = \begin{pmatrix}3&6\\6&14\end{pmatrix}, \mathbf X'\mathbf Y = \begin{pmatrix}11\\25\end{pmatrix}\).

Determinant 6, inverse \((1/6)\begin{pmatrix}14&-6\\-6&3\end{pmatrix}\).

\(\hat{\boldsymbol\beta} = (1/6)\begin{pmatrix}14&-6\\-6&3\end{pmatrix}\begin{pmatrix}11\\25\end{pmatrix} = (1/6)\begin{pmatrix}4\\9\end{pmatrix} = \begin{pmatrix}0.667\\1.500\end{pmatrix}.\)

So \(\hat Y = 0.667 + 1.5 X\) — confirming the matrix algebra mirrors the scalar formulas.

6. Dummy (Indicator) Variables

A dummy variable takes only the values 0 and 1 to bring a qualitative factor (sex, region, season, before/after a policy) into a regression. For a two-category factor,

\[ Y_i = \beta_0 + \beta_1 X_i + \beta_2 D_i + u_i, \qquad D_i = \begin{cases}1 & \text{category present}\\ 0 & \text{otherwise.}\end{cases} \]

Here \(\beta_2\) is the difference in intercept between the two groups (a shift in level); the group with \(D = 0\) is the base (reference) category. A factor with \(m\) categories needs \(m - 1\) dummies — including all \(m\) plus an intercept causes perfect collinearity, the dummy-variable trap. Interacting a dummy with \(X\) (a term \(D_i X_i\)) lets the slope differ between groups as well.

EXAMPLE

\(\widehat{\text{Wage}} = 8 + 1.5\,\text{Educ} + 3\,D\) with \(D = 1\) for urban workers. Urban workers earn, on average, 3 units more than rural workers with the same education; rural is the base category.

Key Take-aways from Unit 2