Multicollinearity means a linear relationship exists among two or more regressors in the model. It is a data problem, not a violation of a CLRM assumption.
with at least one \(\lambda_j \ne 0\). One regressor is an exact linear combination of others. Then \((\mathbf X'\mathbf X)\) is singular, the inverse does not exist, and OLS cannot be computed — the model is unidentified.
The regressors are highly but not perfectly correlated. \((\mathbf X'\mathbf X)\) is invertible but nearly singular — the determinant is close to zero. This is the practically common case, and what is meant when economists refer to "multicollinearity."
Under near multicollinearity (the only practically relevant case):
The simplest diagnostic. Suggestive, not definitive.
Compute the correlation matrix. \(|r_{X_j X_k}| > 0.8\) raises a red flag.
where \(R_j^2\) is the coefficient of determination from regressing \(X_j\) on all the other regressors. Rules of thumb:
The reciprocal of VIF. Smaller tolerance = more collinearity. Critical values: TOL \(< 0.2\) is concerning, TOL \(< 0.1\) is severe.
Let \(\lambda_{\max}\) and \(\lambda_{\min}\) be the largest and smallest eigenvalues of \(\mathbf X'\mathbf X\). The condition number is
\[ \kappa = \sqrt{\lambda_{\max}/\lambda_{\min}}. \]Rule of thumb: \(\kappa < 10\) acceptable, \(\kappa\) between 10 and 30 moderate, \(\kappa > 30\) severe.
Regress each \(X_j\) on the others. If any auxiliary \(R_j^2 > R^2\) of the main regression, multicollinearity is "harmful."
In a 3-regressor model, regressing \(X_1\) on \(X_2, X_3\) gives \(R_1^2 = 0.92\). Then:
VIF\(_1\) = \(1/(1 - 0.92) = 12.5\); TOL\(_1\) = 0.08.
VIF \(> 10\) and TOL \(< 0.1\) — severe collinearity of \(X_1\) with the other regressors. Its individual \(t\)-statistic is unreliable.
Main regression: \(\hat\beta_1 = 0.45\), \(\mathrm{se}(\hat\beta_1) = 0.30\), \(t = 1.5\) — not significant at 5%. VIF\(_1\) = 12.5 means the variance is 12.5× higher than it would be if \(X_1\) were orthogonal to the rest. In a world without collinearity the SE would be \(0.30/\sqrt{12.5} = 0.085\), giving \(t = 5.3\) — highly significant. The economic relationship is real; the multicollinearity merely hides it statistically.
If the goal is prediction rather than individual coefficient interpretation, and the collinearity is stable, multicollinearity does not bias predictions. \(\hat Y\) and its standard error are still correct. Sometimes the cheapest fix is to live with it.
If two regressors essentially measure the same thing (e.g. years of education and years of schooling), keeping both adds noise without information. Dropping one removes the problem — but introduces specification bias if both genuinely belong in the model.
Variances of OLS estimators are inversely related to \(n\) and \(S_{jj}\). More data widens the range of regressors (more variation) and may reduce collinearity.
Replace the original regressors with their first few principal components — orthogonal by construction. Loses interpretability but eliminates collinearity.
Add a penalty term \(\lambda \|\boldsymbol\beta\|^2\) to the OLS criterion:
\[ \hat{\boldsymbol\beta}_{\text{ridge}} = (\mathbf X'\mathbf X + \lambda \mathbf I)^{-1}\mathbf X'\mathbf Y. \]This adds a small constant to the diagonal of \(\mathbf X'\mathbf X\), making it well-conditioned even when nearly singular. Introduces a tiny bias in exchange for greatly reduced variance — "biased but stable."
External information from one type of data (e.g. cross-section gives a known price elasticity) can be imposed as a constraint on the other (time-series).
A wage equation includes both years of education and years of schooling — nearly identical concepts in many surveys. VIF for each is > 30. The correct remedy is to drop one (they measure the same construct), not to apply ridge regression. Always start with the economic interpretation.
A demand model with 8 prices that all move together (typical macro time series). Dropping variables would mis-specify the economic model. Ridge regression with \(\lambda = 0.1\) shrinks the coefficient estimates toward zero, dramatically reducing their variance. Cross-validation can pick the optimal \(\lambda\). Coefficients are now individually interpretable, at the cost of a small bias.
| Issue | Property affected | Detection | Remedy |
|---|---|---|---|
| Heteroscedasticity | Efficiency | BP, White | WLS, robust SE |
| Autocorrelation | Efficiency | Durbin–Watson | GLS, Cochrane–Orcutt |
| Multicollinearity | Variance / SE only | VIF, TOL | Drop, ridge, more data |
| Specification error | Bias | RESET | Re-specify |