The general linear model (GLM) is
\[ \mathbf{y} = X\boldsymbol\beta + \boldsymbol\varepsilon, \qquad E(\boldsymbol\varepsilon) = \mathbf{0}, \qquad \operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2} I_n, \]with \(\mathbf{y}\) the \(n \times 1\) observation vector, \(X\) the \(n \times p\) design matrix of known constants, \(\boldsymbol\beta\) the \(p \times 1\) parameter vector and \(\boldsymbol\varepsilon\) the error vector. Together the last two conditions are the Gauss–Markov assumptions: zero mean, constant variance, zero covariance. Note what is not assumed — normality is not needed for anything in this unit.
What "general" buys. One formulation covers regression, analysis of variance and analysis of covariance; only \(X\) changes.
| model | a row of \(X\) | rank |
|---|---|---|
| simple regression | \((1, x_i)\) | 2, full |
| multiple regression, \(k\) regressors | \((1, x_{i1}, \ldots, x_{ik})\) | \(k+1\), full if no exact collinearity |
| one-way layout, \(k\) treatments | \((1, 0, \ldots, 1, \ldots, 0)\) | \(k\), but \(p = k + 1\) columns — not full |
The last row is the case the Foundation treatment cannot handle, and it is the reason the rest of this unit exists.
A linear parametric function \(\boldsymbol\ell'\boldsymbol\beta\) is estimable if there exists a vector \(\mathbf{a}\) such that
\[ E\left(\mathbf{a}'\mathbf{y}\right) = \boldsymbol\ell'\boldsymbol\beta \quad \text{for every } \boldsymbol\beta, \]that is, some linear function of the observations is unbiased for it.
The criterion. Since \(E(\mathbf{a}'\mathbf{y}) = \mathbf{a}'X\boldsymbol\beta\), the requirement is \(\mathbf{a}'X\boldsymbol\beta = \boldsymbol\ell'\boldsymbol\beta\) for all \(\boldsymbol\beta\), which holds if and only if
\[ \boldsymbol\ell' = \mathbf{a}'X \quad \text{for some } \mathbf{a}, \]i.e. \(\boldsymbol\ell'\) lies in the row space of \(X\). Two equivalent working tests:
\[ \operatorname{rank}\begin{pmatrix} X \\ \boldsymbol\ell' \end{pmatrix} = \operatorname{rank}(X), \qquad\text{or}\qquad \boldsymbol\ell'\left(X'X\right)^{-}X'X = \boldsymbol\ell'. \]When \(X\) has full column rank every \(\boldsymbol\ell\) qualifies, and the question does not arise. It arises exactly when it does not.
Given. Two treatments, two observations each, under the over-parameterised model \(y_{ij} = \mu + \tau_i + \varepsilon_{ij}\) with \(\boldsymbol\beta = (\mu, \tau_1, \tau_2)'\):
\[ X = \begin{pmatrix} 1 & 1 & 0 \\ 1 & 1 & 0 \\ 1 & 0 & 1 \\ 1 & 0 & 1 \end{pmatrix}. \]Step 1 — the rank. Column 1 is the sum of columns 2 and 3, so the three columns are dependent and \(\operatorname{rank}(X) = 2 < 3 = p\). Consequently
\[ \det\left(X'X\right) = 0, \]and \((X'X)^{-1}\) does not exist. There are infinitely many solutions to the normal equations.
Step 2 — the row space. \(X\) has only two distinct rows, \((1,1,0)\) and \((1,0,1)\), so the row space is
\[ \left\{ a(1,1,0) + b(1,0,1) \right\} = \left\{ (a+b,\ a,\ b) : a, b \in \mathbb{R} \right\}. \]Step 3 — test each candidate by solving for \(a\) and \(b\).
| function | \(\boldsymbol\ell'\) | solve \((a+b, a, b) = \boldsymbol\ell'\) | estimable? |
|---|---|---|---|
| \(\mu\) | \((1,0,0)\) | \(a = 0\), \(b = 0\), but then \(a+b = 0 \ne 1\) | no |
| \(\tau_1\) | \((0,1,0)\) | \(a = 1\), \(b = 0\), but then \(a+b = 1 \ne 0\) | no |
| \(\mu + \tau_1\) | \((1,1,0)\) | \(a = 1\), \(b = 0\), and \(a+b = 1\) \(\checkmark\) | yes |
| \(\tau_1 - \tau_2\) | \((0,1,-1)\) | \(a = 1\), \(b = -1\), and \(a+b = 0\) \(\checkmark\) | yes |
Step 4 — read the pattern. The estimable functions are exactly those of the form \(c_0\mu + c_1\tau_1 + c_2\tau_2\) with \(c_0 = c_1 + c_2\). Treatment contrasts, where \(c_0 = 0\) and \(\sum_i c_i = 0\), always qualify.
Step 5 — estimate what is estimable. With observations \(12, 14\) on treatment 1 and \(18, 20\) on treatment 2,
\[ \bar y_1 = \frac{12 + 14}{2} = 13, \qquad \bar y_2 = \frac{18 + 20}{2} = 19, \] \[ \widehat{\tau_1 - \tau_2} = \bar y_1 - \bar y_2 = 13 - 19 = -6. \]Step 6 — the error sum of squares and its degrees of freedom.
\[ SSE = (12-13)^{2} + (14-13)^{2} + (18-19)^{2} + (20-19)^{2} = 1 + 1 + 1 + 1 = 4, \] \[ \text{d.f.} = n - \operatorname{rank}(X) = 4 - 2 = 2, \qquad s^{2} = \frac{4}{2} = 2. \]Note the degrees of freedom use the rank, not the number of columns. Using \(4 - 3 = 1\) would be wrong and would inflate \(s^{2}\) by a factor of two.
Step 7 — a test on the estimable contrast.
\[ \operatorname{Var}\left(\bar y_1 - \bar y_2\right) = \sigma^{2}\left(\tfrac12 + \tfrac12\right) = \sigma^{2}, \qquad \widehat{SE} = \sqrt{s^{2}} = \sqrt{2} = 1.414214, \] \[ t = \frac{-6}{1.414214} = -4.242641 \quad \text{on } 2 \text{ d.f.}, \qquad P\left(|t_2| > 4.242641\right) = 0.051317. \]Interpretation. A six-unit difference with only two error degrees of freedom falls just short of significance at \(5\%\). Note that no amount of algebra could have produced a test of \(\tau_1\) alone: it is not estimable, so there is nothing to test. Imposing a side condition such as \(\tau_1 + \tau_2 = 0\) makes \(\tau_1\) computable, but what is then computed depends on the side condition chosen, whereas the contrast \(\tau_1 - \tau_2\) does not. Estimable functions are the ones the data actually determine.
Under the Gauss–Markov assumptions, for any estimable \(\boldsymbol\ell'\boldsymbol\beta\), the least squares estimator \(\boldsymbol\ell'\hat{\boldsymbol\beta}\) is the Best Linear Unbiased Estimator (BLUE): among all linear unbiased estimators it has the smallest variance.
"Best" means minimum variance; "linear" means linear in \(\mathbf{y}\); "unbiased" means correct on average. Normality is not required — only the three assumptions on the first two moments.
Step 1 — set up the competition. Let \(\mathbf{a}_0'\mathbf{y}\) be the least squares estimator of \(\boldsymbol\ell'\boldsymbol\beta\), and let \(\mathbf{a}'\mathbf{y}\) be any other linear unbiased estimator of the same thing. Write
\[ \mathbf{a} = \mathbf{a}_0 + \mathbf{d}, \qquad \mathbf{d} = \mathbf{a} - \mathbf{a}_0. \]Step 2 — unbiasedness forces \(\mathbf{d}\) into a particular space. Both are unbiased, so
\[ \mathbf{a}'X\boldsymbol\beta = \boldsymbol\ell'\boldsymbol\beta = \mathbf{a}_0'X\boldsymbol\beta \quad \text{for every } \boldsymbol\beta, \]hence \((\mathbf{a} - \mathbf{a}_0)'X = \mathbf{0}'\), that is \(\mathbf{d}'X = \mathbf{0}'\). So \(\mathbf{d}\) is orthogonal to every column of \(X\), i.e. \(\mathbf{d}\) lies in the error space.
Step 3 — \(\mathbf{a}_0\) lies in the estimation space. The least squares estimator is a function of the fitted values, which lie in \(\mathcal{C}(X)\), so \(\mathbf{a}_0 = X\mathbf{c}\) for some \(\mathbf{c}\).
Step 4 — the cross term vanishes.
\[ \mathbf{a}_0'\mathbf{d} = \left(X\mathbf{c}\right)'\mathbf{d} = \mathbf{c}'X'\mathbf{d} = \mathbf{c}'\left(\mathbf{d}'X\right)' = \mathbf{c}'\mathbf{0} = 0, \]using Step 2.
Step 5 — expand the variance. Since \(\operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2}I\), \(\operatorname{Var}(\mathbf{a}'\mathbf{y}) = \sigma^{2}\mathbf{a}'\mathbf{a}\), so
\[ \operatorname{Var}\left(\mathbf{a}'\mathbf{y}\right) = \sigma^{2}\left(\mathbf{a}_0 + \mathbf{d}\right)'\left(\mathbf{a}_0 + \mathbf{d}\right) = \sigma^{2}\left(\mathbf{a}_0'\mathbf{a}_0 + 2\,\mathbf{a}_0'\mathbf{d} + \mathbf{d}'\mathbf{d}\right). \]Step 6 — drop the cross term and conclude. By Step 4 the middle term is zero, so
\[ \operatorname{Var}\left(\mathbf{a}'\mathbf{y}\right) = \sigma^{2}\,\mathbf{a}_0'\mathbf{a}_0 + \sigma^{2}\,\mathbf{d}'\mathbf{d} = \operatorname{Var}\left(\mathbf{a}_0'\mathbf{y}\right) + \sigma^{2}\|\mathbf{d}\|^{2}. \]The added term is non-negative and is zero only when \(\mathbf{d} = \mathbf{0}\), that is when the competitor is the least squares estimator. \(\blacksquare\)
A linear zero function is \(\mathbf{d}'\mathbf{y}\) with \(E(\mathbf{d}'\mathbf{y}) = 0\) for every \(\boldsymbol\beta\) — equivalently \(\mathbf{d}'X = \mathbf{0}'\), which is exactly the condition met in Step 2 above. Step 4 then says:
\[ \textbf{an estimator is BLUE if and only if it is uncorrelated with every linear zero function.} \]This is the cleanest characterisation of least squares. \(\mathbb{R}^{n}\) splits into the estimation space \(\mathcal{C}(X)\), which carries all the information about \(\boldsymbol\beta\), and the error space, its orthogonal complement, which carries none. Any estimator that borrows from the error space adds variance and no information.
Suppose the third assumption is replaced by
\[ \operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2}V, \qquad V \text{ known, positive definite}, \]which covers heteroscedasticity (\(V\) diagonal with unequal entries) and autocorrelation (\(V\) with non-zero off-diagonals), the two failures diagnosed in Econometrics, Unit 3 and Unit 5.
Aitken's estimator.
\[ \hat{\boldsymbol\beta}_{GLS} = \left(X'V^{-1}X\right)^{-1}X'V^{-1}\mathbf{y}, \qquad \operatorname{Var}\left(\hat{\boldsymbol\beta}_{GLS}\right) = \sigma^{2}\left(X'V^{-1}X\right)^{-1}, \]and it is BLUE under the modified assumptions.
Step 1. \(V\) is symmetric positive definite, so by the spectral decomposition of Unit 2 it has a symmetric positive definite square root, and \(V^{-1/2}\) exists.
Step 2. Premultiply the whole model by \(V^{-1/2}\):
\[ V^{-1/2}\mathbf{y} = V^{-1/2}X\boldsymbol\beta + V^{-1/2}\boldsymbol\varepsilon, \quad\text{say}\quad \mathbf{y}^{*} = X^{*}\boldsymbol\beta + \boldsymbol\varepsilon^{*}. \]Step 3. The transformed errors satisfy the original assumptions:
\[ \operatorname{Var}\left(\boldsymbol\varepsilon^{*}\right) = V^{-1/2}\left(\sigma^{2}V\right)V^{-1/2} = \sigma^{2} V^{-1/2}V^{1/2}V^{1/2}V^{-1/2} = \sigma^{2}I. \]Step 4. So ordinary least squares applies to the starred model, and Gauss–Markov applies to it unchanged:
\[ \hat{\boldsymbol\beta} = \left(X^{*\prime}X^{*}\right)^{-1}X^{*\prime}\mathbf{y}^{*} = \left(X'V^{-1}X\right)^{-1}X'V^{-1}\mathbf{y}, \]since \(X^{*\prime}X^{*} = X'V^{-1/2}V^{-1/2}X = X'V^{-1}X\). \(\blacksquare\)
The cost of ignoring \(V\). Ordinary least squares stays unbiased when \(V \ne I\) — unbiasedness needs only \(E(\boldsymbol\varepsilon) = \mathbf{0}\). What it loses is efficiency, and, more damagingly, the reported standard error \(s^{2}(X'X)^{-1}\) is then the wrong formula, so every \(t\) and \(F\) is wrong. The coefficients are believable; the inference is not.
Exact multicollinearity means \(\operatorname{rank}(X) < p\): some column is an exact linear combination of others, \(X'X\) is singular, and the parameters are not all estimable — the situation of Example 4.1.
Near multicollinearity means \(X'X\) is non-singular but close to singular. Everything is technically estimable and nothing is usefully estimated. The symptoms:
Diagnosis. The variance inflation factor for regressor \(j\) is
\[ VIF_j = \frac{1}{1 - R_j^{2}}, \]with \(R_j^{2}\) the coefficient of determination from regressing \(x_j\) on all the other regressors. It is the factor by which \(\operatorname{Var}(\hat\beta_j)\) is multiplied, relative to what it would be with orthogonal regressors. \(VIF > 10\) — that is \(R_j^{2} > 0.9\) — is the usual warning line.
Given. Three regressors observed on five units:
| unit | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| \(x_1\) | 1 | 2 | 3 | 4 | 5 |
| \(x_2\) | 2 | 4 | 6 | 8 | 11 |
| \(x_3\) | 1 | 3 | 2 | 5 | 4 |
Step 1 — look before computing. \(x_2\) is exactly \(2x_1\) at the first four units and departs only at the fifth, where it is \(11\) rather than \(10\). So the collinearity is near-exact but not exact, and \(X'X\) is non-singular — just barely.
Step 2 — the pairwise correlations.
\[ r(x_1, x_2) = 0.995893, \qquad r(x_1, x_3) = 0.800000, \qquad r(x_2, x_3) = 0.769554. \]Step 3 — the variance inflation factors. Each comes from regressing one column on the other two:
| regressor | \(R_j^{2}\) | \(VIF_j = 1/(1 - R_j^{2})\) | reading |
|---|---|---|---|
| \(x_1\) | 0.99457286 | 184.2593 | severe |
| \(x_2\) | 0.99385246 | 162.6667 | severe |
| \(x_3\) | 0.73000000 | 3.7037 | acceptable |
Step 4 — read what the numbers say. The variance of \(\hat\beta_1\) is \(184\) times larger than it would be with orthogonal regressors, so its standard error is \(\sqrt{184.2593} = 13.57\) times larger. A coefficient would have to be enormous to clear that standard error. Meanwhile \(x_3\), with \(VIF = 3.70\), is not part of the problem at all.
Step 5 — note what is not wrong. The estimates remain unbiased and the model's predictions remain good inside the range of the data. Multicollinearity damages the separation of effects, not the fit. That is why dropping one of \(x_1, x_2\) usually costs almost nothing in \(R^{2}\) while transforming the standard errors.
Interpretation. The remedies follow from the diagnosis: drop one of the pair, combine them into a single index, collect data that breaks the pattern, or accept a little bias for a large variance reduction by using ridge regression. What does not help is a larger sample of the same shape — the problem is in the design, not the sample size.
| Application | What \(X\) holds | Where on this site |
|---|---|---|
| Simple and multiple regression | a column of ones and the regressors | Statistical Methods, Unit 4 |
| Analysis of variance | indicator columns, rank deficient | Design and Analysis of Experiments |
| Analysis of covariance | indicators and continuous columns together | Design and Analysis of Experiments (STS-203) |
| Polynomial and response surface models | powers and cross-products of the inputs | Design and Analysis of Experiments (STS-203) |
| Weighted least squares | \(V\) diagonal with unequal variances | Econometrics, Unit 3 |
| Time series with autocorrelated errors | \(V\) with an AR(1) structure | Econometrics, Unit 5 |
Each of these is the same estimator, the same theorem and the same geometry, with a different \(X\) and a different \(V\). That economy is the reason the general form is worth learning once rather than learning six special cases.