The population regression function (PRF):
\[ E(Y \mid X_i) = \beta_0 + \beta_1 X_i. \]Each individual observation is
\[ Y_i = \beta_0 + \beta_1 X_i + u_i, \]where \(\beta_0\) is the intercept, \(\beta_1\) is the slope, and \(u_i\) is the stochastic disturbance.
OLS chooses \(\hat\beta_0, \hat\beta_1\) to minimise the residual sum of squares:
\[ \min_{\hat\beta_0,\hat\beta_1}\ \sum_{i=1}^{n} \hat u_i^2 = \sum_{i=1}^{n}(Y_i - \hat\beta_0 - \hat\beta_1 X_i)^2. \]Setting partial derivatives to zero:
\[ \sum Y_i = n\hat\beta_0 + \hat\beta_1\sum X_i, \qquad \sum X_i Y_i = \hat\beta_0\sum X_i + \hat\beta_1\sum X_i^2. \]Solving:
\[ \hat\beta_1 = \frac{\sum (X_i - \bar X)(Y_i - \bar Y)}{\sum (X_i - \bar X)^2} = \frac{S_{XY}}{S_{XX}}, \] \[ \hat\beta_0 = \bar Y - \hat\beta_1 \bar X. \]and \(\sigma^2\) is estimated by \(\hat\sigma^2 = \sum \hat u_i^2 /(n-2)\), an unbiased estimator.
Data on 5 households (\(X\) = income, \(Y\) = consumption, ₹'000):
| i | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| \(X_i\) | 10 | 20 | 30 | 40 | 50 |
| \(Y_i\) | 11 | 20 | 27 | 34 | 43 |
\(\bar X = 30, \bar Y = 27\), \(S_{XY} = \sum(X_i-30)(Y_i-27) = 780\), \(S_{XX} = \sum(X_i-30)^2 = 1000\).
\(\hat\beta_1 = 780/1000 = 0.78\), \(\hat\beta_0 = 27 - 0.78\times 30 = 3.60\).
Fitted line: \(\hat Y = 3.60 + 0.78 X\). Marginal propensity to consume is 0.78.
From Example 1, predict consumption for a household with income ₹35,000: \(\hat Y = 3.60 + 0.78 \times 35 = 30.90\) (i.e. ₹30,900).
For income ₹100,000 (extrapolation beyond data): \(\hat Y = 3.60 + 0.78 \times 100 = 81.60\) (₹81,600). Note: extrapolation is dangerous — the linear relation may not hold outside the sampled range.
Under the classical assumptions 1–8 (linearity, exogeneity, homoscedasticity, no serial correlation, etc.), the OLS estimators \(\hat\beta_0, \hat\beta_1\) are the Best Linear Unbiased Estimators (BLUE) of the population parameters \(\beta_0, \beta_1\).
Normality of \(u\) is not required for the Gauss–Markov theorem — only for distribution-based inference (\(t\), \(F\) tests).
When CLRM assumptions hold, you cannot find a "better" linear unbiased estimator than OLS. Therefore most diagnostics in this course test whether one of the assumptions is violated — and if so, which alternative estimator (GLS, WLS, FGLS, robust SE, etc.) restores BLUE-ness.
For three variables \(Y, X_1, X_2\):
\[ R^2_{Y \cdot X_1 X_2} = \frac{r^2_{YX_1} + r^2_{YX_2} - 2\,r_{YX_1}r_{YX_2}r_{X_1 X_2}}{1 - r^2_{X_1 X_2}}. \]\(R\) measures the linear association of \(Y\) with the joint influence of \(X_1\) and \(X_2\). It is the correlation between \(Y\) and \(\hat Y\) from the fitted regression.
Partial correlation between \(Y\) and \(X_1\), holding \(X_2\) constant:
\[ r_{YX_1 \cdot X_2} = \frac{r_{YX_1} - r_{YX_2}r_{X_1 X_2}}{\sqrt{(1 - r^2_{YX_2})(1 - r^2_{X_1 X_2})}}. \]It removes the influence of \(X_2\) on both \(Y\) and \(X_1\) before measuring their association — the "pure" effect of \(X_1\) on \(Y\).
\(r_{YX_1} = 0.80, r_{YX_2} = 0.60, r_{X_1 X_2} = 0.50\). Find \(r_{YX_1 \cdot X_2}\).
\(r_{YX_1 \cdot X_2} = (0.80 - 0.60 \times 0.50)/\sqrt{(1-0.36)(1-0.25)} = 0.50/\sqrt{0.48} = 0.50/0.6928 = 0.7217.\)
Net of \(X_2\), \(Y\) and \(X_1\) are still strongly positively associated.
Using the same numbers:
\(R^2_{Y \cdot X_1 X_2} = (0.64 + 0.36 - 2 \times 0.80 \times 0.60 \times 0.50)/(1 - 0.25) = (1.00 - 0.48)/0.75 = 0.6933.\)
\(R = 0.8326\). Together \(X_1\) and \(X_2\) explain about 69% of the variation in \(Y\).
With \(k\) explanatory variables:
\[ Y_i = \beta_0 + \beta_1 X_{1i} + \beta_2 X_{2i} + \cdots + \beta_k X_{ki} + u_i, \quad i = 1, \ldots, n. \]In matrix form:
\[ \mathbf Y = \mathbf X \boldsymbol\beta + \mathbf u, \]where \(\mathbf Y\) is \(n \times 1\), \(\mathbf X\) is \(n \times (k+1)\), \(\boldsymbol\beta\) is \((k+1) \times 1\), and \(\mathbf u\) is \(n \times 1\).
This is the single most important formula in the entire course — every multiple-regression derivation springs from it. The variance-covariance matrix is
\[ \mathrm{Var}(\hat{\boldsymbol\beta}) = \sigma^2 (\mathbf X' \mathbf X)^{-1}. \]Mincer wage equation: \(\log W_i = \beta_0 + \beta_1\,\text{educ}_i + \beta_2\,\text{exper}_i + \beta_3\,\text{exper}_i^2 + u_i\). With \(n = 526\) observations from the CPS (US), typical estimates are \(\hat\beta_1 = 0.092\) (each extra year of education raises wages by 9.2%), \(\hat\beta_2 = 0.041\), \(\hat\beta_3 = -0.0007\) (wage rises with experience but at a diminishing rate).
For 3 observations \((X, Y) = (1, 2), (2, 4), (3, 5)\), the design matrix (with intercept) is
\(\mathbf X = \begin{pmatrix}1&1\\1&2\\1&3\end{pmatrix}, \mathbf Y = \begin{pmatrix}2\\4\\5\end{pmatrix}\).
\(\mathbf X'\mathbf X = \begin{pmatrix}3&6\\6&14\end{pmatrix}, \mathbf X'\mathbf Y = \begin{pmatrix}11\\25\end{pmatrix}\).
Determinant 6, inverse \((1/6)\begin{pmatrix}14&-6\\-6&3\end{pmatrix}\).
\(\hat{\boldsymbol\beta} = (1/6)\begin{pmatrix}14&-6\\-6&3\end{pmatrix}\begin{pmatrix}11\\25\end{pmatrix} = (1/6)\begin{pmatrix}4\\9\end{pmatrix} = \begin{pmatrix}0.667\\1.500\end{pmatrix}.\)
So \(\hat Y = 0.667 + 1.5 X\) — confirming the matrix algebra mirrors the scalar formulas.
A dummy variable takes only the values 0 and 1 to bring a qualitative factor (sex, region, season, before/after a policy) into a regression. For a two-category factor,
Here \(\beta_2\) is the difference in intercept between the two groups (a shift in level); the group with \(D = 0\) is the base (reference) category. A factor with \(m\) categories needs \(m - 1\) dummies — including all \(m\) plus an intercept causes perfect collinearity, the dummy-variable trap. Interacting a dummy with \(X\) (a term \(D_i X_i\)) lets the slope differ between groups as well.
\(\widehat{\text{Wage}} = 8 + 1.5\,\text{Educ} + 3\,D\) with \(D = 1\) for urban workers. Urban workers earn, on average, 3 units more than rural workers with the same education; rural is the base category.