Skip to the content

Topics Covered

Concept Perfect vs Imperfect Consequences Detection VIF Tolerance Condition Number Remedies
On this page
  1. 1. Concept of Multicollinearity
  2. 2. Consequences of Multicollinearity
  3. 3. Detection of Multicollinearity
  4. 4. Remedies for Multicollinearity
  5. 5. Multicollinearity vs Other Problems
  6. Key Take-aways from Unit 4

1. Concept of Multicollinearity

DEFINITION

Multicollinearity means a linear relationship exists among two or more regressors in the model. It is a data problem, not a violation of a CLRM assumption.

1.1 Perfect Multicollinearity

\[ \lambda_1 X_1 + \lambda_2 X_2 + \cdots + \lambda_k X_k = 0, \]

with at least one \(\lambda_j \ne 0\). One regressor is an exact linear combination of others. Then \((\mathbf X'\mathbf X)\) is singular, the inverse does not exist, and OLS cannot be computed — the model is unidentified.

1.2 Imperfect (Near) Multicollinearity

The regressors are highly but not perfectly correlated. \((\mathbf X'\mathbf X)\) is invertible but nearly singular — the determinant is close to zero. This is the practically common case, and what is meant when economists refer to "multicollinearity."

1.3 Common Sources

2. Consequences of Multicollinearity

Under near multicollinearity (the only practically relevant case):

  1. OLS estimators remain BLUE — unbiased, consistent, minimum variance among linear unbiased estimators.
  2. But the variance is large: \(\mathrm{Var}(\hat\beta_j) = \sigma^2/(S_{jj}(1-R_j^2))\) where \(R_j^2\) is the \(R^2\) from regressing \(X_j\) on the other regressors. As \(R_j^2 \to 1\), \(\mathrm{Var}(\hat\beta_j) \to \infty\).
  3. Standard errors are inflated, so \(t\)-statistics are small and individual coefficients appear insignificant.
  4. Yet \(R^2\) can be very high and the joint \(F\)-test highly significant — a tell-tale symptom.
  5. OLS estimates become highly sensitive to small changes in the data — adding or dropping one observation can flip a coefficient's sign.
  6. Confidence intervals are wide.
  7. Coefficient signs may be wrong relative to economic theory.
Classic warning sign: High overall fit (\(R^2 = 0.95\), \(F = 250\)), but every individual \(t\)-statistic is below 2. The regressors are jointly powerful but individually indistinguishable — a textbook multicollinearity diagnosis.

3. Detection of Multicollinearity

3.1 High \(R^2\) but Few Significant \(t\)s

The simplest diagnostic. Suggestive, not definitive.

3.2 High Pairwise Correlations Between Regressors

Compute the correlation matrix. \(|r_{X_j X_k}| > 0.8\) raises a red flag.

3.3 Variance Inflation Factor (VIF)

\[ \text{VIF}_j = \frac{1}{1 - R_j^2}, \]

where \(R_j^2\) is the coefficient of determination from regressing \(X_j\) on all the other regressors. Rules of thumb:

3.4 Tolerance (TOL)

\[ \text{TOL}_j = 1 - R_j^2 = \frac{1}{\text{VIF}_j}. \]

The reciprocal of VIF. Smaller tolerance = more collinearity. Critical values: TOL \(< 0.2\) is concerning, TOL \(< 0.1\) is severe.

3.5 Eigenvalues and Condition Number

Let \(\lambda_{\max}\) and \(\lambda_{\min}\) be the largest and smallest eigenvalues of \(\mathbf X'\mathbf X\). The condition number is

\[ \kappa = \sqrt{\lambda_{\max}/\lambda_{\min}}. \]

Rule of thumb: \(\kappa < 10\) acceptable, \(\kappa\) between 10 and 30 moderate, \(\kappa > 30\) severe.

3.6 Auxiliary Regressions (Klein's Rule)

Regress each \(X_j\) on the others. If any auxiliary \(R_j^2 > R^2\) of the main regression, multicollinearity is "harmful."

EXAMPLE 1 — Computing VIF

In a 3-regressor model, regressing \(X_1\) on \(X_2, X_3\) gives \(R_1^2 = 0.92\). Then:

VIF\(_1\) = \(1/(1 - 0.92) = 12.5\); TOL\(_1\) = 0.08.

VIF \(> 10\) and TOL \(< 0.1\) — severe collinearity of \(X_1\) with the other regressors. Its individual \(t\)-statistic is unreliable.

EXAMPLE 2 — Inflated standard errors

Main regression: \(\hat\beta_1 = 0.45\), \(\mathrm{se}(\hat\beta_1) = 0.30\), \(t = 1.5\) — not significant at 5%. VIF\(_1\) = 12.5 means the variance is 12.5× higher than it would be if \(X_1\) were orthogonal to the rest. In a world without collinearity the SE would be \(0.30/\sqrt{12.5} = 0.085\), giving \(t = 5.3\) — highly significant. The economic relationship is real; the multicollinearity merely hides it statistically.

4. Remedies for Multicollinearity

4.1 Do Nothing

If the goal is prediction rather than individual coefficient interpretation, and the collinearity is stable, multicollinearity does not bias predictions. \(\hat Y\) and its standard error are still correct. Sometimes the cheapest fix is to live with it.

4.2 Drop a Redundant Variable

If two regressors essentially measure the same thing (e.g. years of education and years of schooling), keeping both adds noise without information. Dropping one removes the problem — but introduces specification bias if both genuinely belong in the model.

4.3 Increase the Sample Size

Variances of OLS estimators are inversely related to \(n\) and \(S_{jj}\). More data widens the range of regressors (more variation) and may reduce collinearity.

4.4 Transform the Variables

4.5 Use Principal Components Regression (PCR)

Replace the original regressors with their first few principal components — orthogonal by construction. Loses interpretability but eliminates collinearity.

4.6 Ridge Regression

Add a penalty term \(\lambda \|\boldsymbol\beta\|^2\) to the OLS criterion:

\[ \hat{\boldsymbol\beta}_{\text{ridge}} = (\mathbf X'\mathbf X + \lambda \mathbf I)^{-1}\mathbf X'\mathbf Y. \]

This adds a small constant to the diagonal of \(\mathbf X'\mathbf X\), making it well-conditioned even when nearly singular. Introduces a tiny bias in exchange for greatly reduced variance — "biased but stable."

4.7 Pooling Cross-Section and Time-Series

External information from one type of data (e.g. cross-section gives a known price elasticity) can be imposed as a constraint on the other (time-series).

EXAMPLE 1 — Choosing a remedy

A wage equation includes both years of education and years of schooling — nearly identical concepts in many surveys. VIF for each is > 30. The correct remedy is to drop one (they measure the same construct), not to apply ridge regression. Always start with the economic interpretation.

EXAMPLE 2 — Ridge regression rescue

A demand model with 8 prices that all move together (typical macro time series). Dropping variables would mis-specify the economic model. Ridge regression with \(\lambda = 0.1\) shrinks the coefficient estimates toward zero, dramatically reducing their variance. Cross-validation can pick the optimal \(\lambda\). Coefficients are now individually interpretable, at the cost of a small bias.

5. Multicollinearity vs Other Problems

IssueProperty affectedDetectionRemedy
HeteroscedasticityEfficiencyBP, WhiteWLS, robust SE
AutocorrelationEfficiencyDurbin–WatsonGLS, Cochrane–Orcutt
MulticollinearityVariance / SE onlyVIF, TOLDrop, ridge, more data
Specification errorBiasRESETRe-specify

Key Take-aways from Unit 4