Usually we test \(H_0: \beta_j = 0\). Reject if \(|t| > t_{\alpha/2, n-k-1}\) or if \(p\)-value \(< \alpha\). The standard error is the square root of the \(j\)-th diagonal of \(\hat\sigma^2 (\mathbf X'\mathbf X)^{-1}\).
To test \(H_0: \beta_1 = \beta_2 = \cdots = \beta_k = 0\) (all slopes simultaneously zero):
\[ F = \frac{R^2 / k}{(1 - R^2)/(n - k - 1)} \sim F_{k,\,n-k-1}. \]A high \(F\) means at least one regressor is statistically significant.
\(R^2 \in [0, 1]\); it measures the proportion of variance in \(Y\) explained by the regressors.
Penalises the addition of irrelevant regressors. Always \(\bar R^2 \le R^2\); equal when \(k = 0\).
| Source | SS | df | MS | F |
|---|---|---|---|---|
| Regression | ESS | \(k\) | ESS/\(k\) | MSR/MSE |
| Residual | RSS | \(n-k-1\) | RSS/\((n-k-1)\) = \(\hat\sigma^2\) | |
| Total | TSS | \(n-1\) |
From Unit 2 Example 1: \(\hat\beta_1 = 0.78\). Suppose \(\mathrm{se}(\hat\beta_1) = 0.05\), \(n = 5\). Then \(t = 0.78/0.05 = 15.6\) with df = 3. Critical \(t_{0.025, 3} = 3.182\). Since \(15.6 \gg 3.182\), reject \(H_0: \beta_1 = 0\) — the MPC is highly significant.
A regression of monthly expenditure on income and family size with \(n = 100\), \(R^2 = 0.62\), \(k = 2\). The joint \(F\):
\(F = (0.62/2)/((1 - 0.62)/(100 - 3)) = 0.31/0.00392 = 79.1\). Critical \(F_{0.05, 2, 97} \approx 3.09\). Reject \(H_0\) overwhelmingly — both regressors jointly explain expenditure.
Classical assumption: \(\mathrm{Var}(u_i) = \sigma^2\) (homoscedasticity). When this fails — \(\mathrm{Var}(u_i) = \sigma_i^2\) varying across observations — we have heteroscedasticity.
Plot \(\hat u_i^2\) against \(\hat Y_i\) or against each \(X_j\). A fan or megaphone pattern suggests heteroscedasticity; a structureless cloud suggests homoscedasticity.
Regress \(\log \hat u_i^2\) on \(\log X_i\):
\[ \log \hat u_i^2 = \alpha + \beta \log X_i + v_i. \]If \(\beta\) is statistically significant, heteroscedasticity is present.
Regress \(|\hat u_i|\) on various functional forms of \(X_i\):
\[ |\hat u_i| = \alpha + \beta X_i + v_i, \quad |\hat u_i| = \alpha + \beta\sqrt{X_i} + v_i,\ \text{etc.} \]If any \(\beta\) is significant, heteroscedasticity is present.
Does not require any assumption about the form of heteroscedasticity — the most general test.
Sort the data by an \(X\) variable, drop the middle observations, run separate OLS on the high and low subsets, and compute \(F = \text{RSS}_{\text{high}}/\text{RSS}_{\text{low}}\). Reject if \(F\) is large.
If \(\mathrm{Var}(u_i) = \sigma_i^2\) is known up to a constant, divide each observation by \(\sigma_i\):
\[ \frac{Y_i}{\sigma_i} = \beta_0 \frac{1}{\sigma_i} + \beta_1 \frac{X_i}{\sigma_i} + \frac{u_i}{\sigma_i}. \]The transformed model has homoscedastic errors and OLS on the transformed model is BLUE.
Also called White's robust standard errors. Adjust only the SE formula, leaving the OLS point estimates unchanged. Used when the form of heteroscedasticity is unknown.
Regress \(Y\) on \(X_1, X_2\), \(n = 50\), \(R^2 = 0.05\) from the auxiliary regression of \(\hat u^2\) on \(X_1, X_2, X_1^2, X_2^2, X_1 X_2\) (so \(d = 5\)). White statistic: \(nR^2 = 50 \times 0.05 = 2.5\). \(\chi^2_{0.05, 5} = 11.07\). Since \(2.5 < 11.07\) we do not reject homoscedasticity. (Low \(R^2\) in the auxiliary regression means residuals are not systematically related to the regressors.)
Per-firm productivity data showing that the variance of error rises with firm size \(L\) (number of employees). Assume \(\sigma_i^2 = \sigma^2 L_i\). Apply WLS by dividing each variable by \(\sqrt{L_i}\):
\(Y_i/\sqrt{L_i} = \beta_0/\sqrt{L_i} + \beta_1 (X_i/\sqrt{L_i}) + u_i^*\), where the new disturbance has variance \(\sigma^2\). Apply OLS to the transformed model — efficiency is restored.
Heteroscedasticity is sometimes a symptom of a deeper modelling mistake. Major types of specification error:
The Ramsey RESET test (Regression Equation Specification Error Test) is the standard general specification check: add powers of \(\hat Y\) to the regression and test their joint significance.
If only \(Y\) is mismeasured but the measurement error is uncorrelated with the regressors, OLS remains unbiased and consistent; variances are inflated, so SEs are larger.
If a regressor \(X\) is observed with error, the OLS slope is biased towards zero ("attenuation bias"):
Where \(X^*\) is the true value. If the error variance is large relative to the true variance, the bias can be severe.
Find a variable \(Z\) correlated with \(X\) but not with the measurement error or other errors. The IV estimator \(\hat\beta_{IV} = (\mathbf Z'\mathbf X)^{-1}\mathbf Z'\mathbf Y\) is consistent for \(\beta\) even with errors in variables.