Skip to the content

Topics Covered

General Linear Model Estimability Gauss–Markov Theorem BLUE Linear Zero Functions Aitken GLS Multicollinearity & VIF
On this page
  1. 1. The General Linear Model
  2. 2. Estimability of a Linear Parametric Function
  3. 3. The Gauss–Markov Theorem
  4. 4. Aitken's Generalized Least Squares
  5. 5. Multicollinearity
  6. 6. Importance and Applications of the GLM
  7. Key Take-aways
Where this unit starts. The classical linear regression model, ordinary least squares, the meaning of BLUE, the Gauss–Markov assumptions and multicollinearity are all taught at Foundation level in Econometrics, Unit 2 and Unit 4, and the two-variable case is worked in Statistical Methods, Unit 4. Those pages treat the full-rank case, where \((X'X)^{-1}\) exists. This unit does the general case: what happens when it does not, which parametric functions can still be estimated, and what replaces ordinary least squares when the errors are neither equally variable nor uncorrelated.

1. The General Linear Model

FORMULATION

The general linear model (GLM) is

\[ \mathbf{y} = X\boldsymbol\beta + \boldsymbol\varepsilon, \qquad E(\boldsymbol\varepsilon) = \mathbf{0}, \qquad \operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2} I_n, \]

with \(\mathbf{y}\) the \(n \times 1\) observation vector, \(X\) the \(n \times p\) design matrix of known constants, \(\boldsymbol\beta\) the \(p \times 1\) parameter vector and \(\boldsymbol\varepsilon\) the error vector. Together the last two conditions are the Gauss–Markov assumptions: zero mean, constant variance, zero covariance. Note what is not assumed — normality is not needed for anything in this unit.

What "general" buys. One formulation covers regression, analysis of variance and analysis of covariance; only \(X\) changes.

modela row of \(X\)rank
simple regression\((1, x_i)\)2, full
multiple regression, \(k\) regressors\((1, x_{i1}, \ldots, x_{ik})\)\(k+1\), full if no exact collinearity
one-way layout, \(k\) treatments\((1, 0, \ldots, 1, \ldots, 0)\)\(k\), but \(p = k + 1\) columns — not full

The last row is the case the Foundation treatment cannot handle, and it is the reason the rest of this unit exists.

2. Estimability of a Linear Parametric Function

DEFINITION AND CRITERION

A linear parametric function \(\boldsymbol\ell'\boldsymbol\beta\) is estimable if there exists a vector \(\mathbf{a}\) such that

\[ E\left(\mathbf{a}'\mathbf{y}\right) = \boldsymbol\ell'\boldsymbol\beta \quad \text{for every } \boldsymbol\beta, \]

that is, some linear function of the observations is unbiased for it.

The criterion. Since \(E(\mathbf{a}'\mathbf{y}) = \mathbf{a}'X\boldsymbol\beta\), the requirement is \(\mathbf{a}'X\boldsymbol\beta = \boldsymbol\ell'\boldsymbol\beta\) for all \(\boldsymbol\beta\), which holds if and only if

\[ \boldsymbol\ell' = \mathbf{a}'X \quad \text{for some } \mathbf{a}, \]

i.e. \(\boldsymbol\ell'\) lies in the row space of \(X\). Two equivalent working tests:

\[ \operatorname{rank}\begin{pmatrix} X \\ \boldsymbol\ell' \end{pmatrix} = \operatorname{rank}(X), \qquad\text{or}\qquad \boldsymbol\ell'\left(X'X\right)^{-}X'X = \boldsymbol\ell'. \]

When \(X\) has full column rank every \(\boldsymbol\ell\) qualifies, and the question does not arise. It arises exactly when it does not.

EXAMPLE 4.1 — WHAT IS ESTIMABLE IN A ONE-WAY LAYOUT

Given. Two treatments, two observations each, under the over-parameterised model \(y_{ij} = \mu + \tau_i + \varepsilon_{ij}\) with \(\boldsymbol\beta = (\mu, \tau_1, \tau_2)'\):

\[ X = \begin{pmatrix} 1 & 1 & 0 \\ 1 & 1 & 0 \\ 1 & 0 & 1 \\ 1 & 0 & 1 \end{pmatrix}. \]

Step 1 — the rank. Column 1 is the sum of columns 2 and 3, so the three columns are dependent and \(\operatorname{rank}(X) = 2 < 3 = p\). Consequently

\[ \det\left(X'X\right) = 0, \]

and \((X'X)^{-1}\) does not exist. There are infinitely many solutions to the normal equations.

Step 2 — the row space. \(X\) has only two distinct rows, \((1,1,0)\) and \((1,0,1)\), so the row space is

\[ \left\{ a(1,1,0) + b(1,0,1) \right\} = \left\{ (a+b,\ a,\ b) : a, b \in \mathbb{R} \right\}. \]

Step 3 — test each candidate by solving for \(a\) and \(b\).

function\(\boldsymbol\ell'\)solve \((a+b, a, b) = \boldsymbol\ell'\)estimable?
\(\mu\)\((1,0,0)\)\(a = 0\), \(b = 0\), but then \(a+b = 0 \ne 1\)no
\(\tau_1\)\((0,1,0)\)\(a = 1\), \(b = 0\), but then \(a+b = 1 \ne 0\)no
\(\mu + \tau_1\)\((1,1,0)\)\(a = 1\), \(b = 0\), and \(a+b = 1\) \(\checkmark\)yes
\(\tau_1 - \tau_2\)\((0,1,-1)\)\(a = 1\), \(b = -1\), and \(a+b = 0\) \(\checkmark\)yes

Step 4 — read the pattern. The estimable functions are exactly those of the form \(c_0\mu + c_1\tau_1 + c_2\tau_2\) with \(c_0 = c_1 + c_2\). Treatment contrasts, where \(c_0 = 0\) and \(\sum_i c_i = 0\), always qualify.

Step 5 — estimate what is estimable. With observations \(12, 14\) on treatment 1 and \(18, 20\) on treatment 2,

\[ \bar y_1 = \frac{12 + 14}{2} = 13, \qquad \bar y_2 = \frac{18 + 20}{2} = 19, \] \[ \widehat{\tau_1 - \tau_2} = \bar y_1 - \bar y_2 = 13 - 19 = -6. \]

Step 6 — the error sum of squares and its degrees of freedom.

\[ SSE = (12-13)^{2} + (14-13)^{2} + (18-19)^{2} + (20-19)^{2} = 1 + 1 + 1 + 1 = 4, \] \[ \text{d.f.} = n - \operatorname{rank}(X) = 4 - 2 = 2, \qquad s^{2} = \frac{4}{2} = 2. \]

Note the degrees of freedom use the rank, not the number of columns. Using \(4 - 3 = 1\) would be wrong and would inflate \(s^{2}\) by a factor of two.

Step 7 — a test on the estimable contrast.

\[ \operatorname{Var}\left(\bar y_1 - \bar y_2\right) = \sigma^{2}\left(\tfrac12 + \tfrac12\right) = \sigma^{2}, \qquad \widehat{SE} = \sqrt{s^{2}} = \sqrt{2} = 1.414214, \] \[ t = \frac{-6}{1.414214} = -4.242641 \quad \text{on } 2 \text{ d.f.}, \qquad P\left(|t_2| > 4.242641\right) = 0.051317. \]

Interpretation. A six-unit difference with only two error degrees of freedom falls just short of significance at \(5\%\). Note that no amount of algebra could have produced a test of \(\tau_1\) alone: it is not estimable, so there is nothing to test. Imposing a side condition such as \(\tau_1 + \tau_2 = 0\) makes \(\tau_1\) computable, but what is then computed depends on the side condition chosen, whereas the contrast \(\tau_1 - \tau_2\) does not. Estimable functions are the ones the data actually determine.

Least squares is one orthogonal projection C(X), the column space — dimension = rank(X) y ŷ = Py = Xb e = (I − P)y X′e = 0 ‖y‖² = ‖ŷ‖² + ‖e‖² — the Pythagorean identity is the ANOVA decomposition schematic: a 3-dimensional picture of an n-dimensional fact, not a plotted dataset
Fig 4.1 — Fitting by least squares drops a perpendicular from y onto the column space of X. The right-angle mark is the normal equations.

3. The Gauss–Markov Theorem

STATEMENT

Under the Gauss–Markov assumptions, for any estimable \(\boldsymbol\ell'\boldsymbol\beta\), the least squares estimator \(\boldsymbol\ell'\hat{\boldsymbol\beta}\) is the Best Linear Unbiased Estimator (BLUE): among all linear unbiased estimators it has the smallest variance.

"Best" means minimum variance; "linear" means linear in \(\mathbf{y}\); "unbiased" means correct on average. Normality is not required — only the three assumptions on the first two moments.

PROOF

Step 1 — set up the competition. Let \(\mathbf{a}_0'\mathbf{y}\) be the least squares estimator of \(\boldsymbol\ell'\boldsymbol\beta\), and let \(\mathbf{a}'\mathbf{y}\) be any other linear unbiased estimator of the same thing. Write

\[ \mathbf{a} = \mathbf{a}_0 + \mathbf{d}, \qquad \mathbf{d} = \mathbf{a} - \mathbf{a}_0. \]

Step 2 — unbiasedness forces \(\mathbf{d}\) into a particular space. Both are unbiased, so

\[ \mathbf{a}'X\boldsymbol\beta = \boldsymbol\ell'\boldsymbol\beta = \mathbf{a}_0'X\boldsymbol\beta \quad \text{for every } \boldsymbol\beta, \]

hence \((\mathbf{a} - \mathbf{a}_0)'X = \mathbf{0}'\), that is \(\mathbf{d}'X = \mathbf{0}'\). So \(\mathbf{d}\) is orthogonal to every column of \(X\), i.e. \(\mathbf{d}\) lies in the error space.

Step 3 — \(\mathbf{a}_0\) lies in the estimation space. The least squares estimator is a function of the fitted values, which lie in \(\mathcal{C}(X)\), so \(\mathbf{a}_0 = X\mathbf{c}\) for some \(\mathbf{c}\).

Step 4 — the cross term vanishes.

\[ \mathbf{a}_0'\mathbf{d} = \left(X\mathbf{c}\right)'\mathbf{d} = \mathbf{c}'X'\mathbf{d} = \mathbf{c}'\left(\mathbf{d}'X\right)' = \mathbf{c}'\mathbf{0} = 0, \]

using Step 2.

Step 5 — expand the variance. Since \(\operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2}I\), \(\operatorname{Var}(\mathbf{a}'\mathbf{y}) = \sigma^{2}\mathbf{a}'\mathbf{a}\), so

\[ \operatorname{Var}\left(\mathbf{a}'\mathbf{y}\right) = \sigma^{2}\left(\mathbf{a}_0 + \mathbf{d}\right)'\left(\mathbf{a}_0 + \mathbf{d}\right) = \sigma^{2}\left(\mathbf{a}_0'\mathbf{a}_0 + 2\,\mathbf{a}_0'\mathbf{d} + \mathbf{d}'\mathbf{d}\right). \]

Step 6 — drop the cross term and conclude. By Step 4 the middle term is zero, so

\[ \operatorname{Var}\left(\mathbf{a}'\mathbf{y}\right) = \sigma^{2}\,\mathbf{a}_0'\mathbf{a}_0 + \sigma^{2}\,\mathbf{d}'\mathbf{d} = \operatorname{Var}\left(\mathbf{a}_0'\mathbf{y}\right) + \sigma^{2}\|\mathbf{d}\|^{2}. \]

The added term is non-negative and is zero only when \(\mathbf{d} = \mathbf{0}\), that is when the competitor is the least squares estimator. \(\blacksquare\)

BLUE'S AND LINEAR ZERO FUNCTIONS

A linear zero function is \(\mathbf{d}'\mathbf{y}\) with \(E(\mathbf{d}'\mathbf{y}) = 0\) for every \(\boldsymbol\beta\) — equivalently \(\mathbf{d}'X = \mathbf{0}'\), which is exactly the condition met in Step 2 above. Step 4 then says:

\[ \textbf{an estimator is BLUE if and only if it is uncorrelated with every linear zero function.} \]

This is the cleanest characterisation of least squares. \(\mathbb{R}^{n}\) splits into the estimation space \(\mathcal{C}(X)\), which carries all the information about \(\boldsymbol\beta\), and the error space, its orthogonal complement, which carries none. Any estimator that borrows from the error space adds variance and no information.

4. Aitken's Generalized Least Squares

WHEN THE ERRORS ARE NOT SPHERICAL

Suppose the third assumption is replaced by

\[ \operatorname{Var}(\boldsymbol\varepsilon) = \sigma^{2}V, \qquad V \text{ known, positive definite}, \]

which covers heteroscedasticity (\(V\) diagonal with unequal entries) and autocorrelation (\(V\) with non-zero off-diagonals), the two failures diagnosed in Econometrics, Unit 3 and Unit 5.

Aitken's estimator.

\[ \hat{\boldsymbol\beta}_{GLS} = \left(X'V^{-1}X\right)^{-1}X'V^{-1}\mathbf{y}, \qquad \operatorname{Var}\left(\hat{\boldsymbol\beta}_{GLS}\right) = \sigma^{2}\left(X'V^{-1}X\right)^{-1}, \]

and it is BLUE under the modified assumptions.

WHY IT WORKS: ONE TRANSFORMATION

Step 1. \(V\) is symmetric positive definite, so by the spectral decomposition of Unit 2 it has a symmetric positive definite square root, and \(V^{-1/2}\) exists.

Step 2. Premultiply the whole model by \(V^{-1/2}\):

\[ V^{-1/2}\mathbf{y} = V^{-1/2}X\boldsymbol\beta + V^{-1/2}\boldsymbol\varepsilon, \quad\text{say}\quad \mathbf{y}^{*} = X^{*}\boldsymbol\beta + \boldsymbol\varepsilon^{*}. \]

Step 3. The transformed errors satisfy the original assumptions:

\[ \operatorname{Var}\left(\boldsymbol\varepsilon^{*}\right) = V^{-1/2}\left(\sigma^{2}V\right)V^{-1/2} = \sigma^{2} V^{-1/2}V^{1/2}V^{1/2}V^{-1/2} = \sigma^{2}I. \]

Step 4. So ordinary least squares applies to the starred model, and Gauss–Markov applies to it unchanged:

\[ \hat{\boldsymbol\beta} = \left(X^{*\prime}X^{*}\right)^{-1}X^{*\prime}\mathbf{y}^{*} = \left(X'V^{-1}X\right)^{-1}X'V^{-1}\mathbf{y}, \]

since \(X^{*\prime}X^{*} = X'V^{-1/2}V^{-1/2}X = X'V^{-1}X\). \(\blacksquare\)

The cost of ignoring \(V\). Ordinary least squares stays unbiased when \(V \ne I\) — unbiasedness needs only \(E(\boldsymbol\varepsilon) = \mathbf{0}\). What it loses is efficiency, and, more damagingly, the reported standard error \(s^{2}(X'X)^{-1}\) is then the wrong formula, so every \(t\) and \(F\) is wrong. The coefficients are believable; the inference is not.

5. Multicollinearity

EXACT AND NEAR COLLINEARITY

Exact multicollinearity means \(\operatorname{rank}(X) < p\): some column is an exact linear combination of others, \(X'X\) is singular, and the parameters are not all estimable — the situation of Example 4.1.

Near multicollinearity means \(X'X\) is non-singular but close to singular. Everything is technically estimable and nothing is usefully estimated. The symptoms:

Diagnosis. The variance inflation factor for regressor \(j\) is

\[ VIF_j = \frac{1}{1 - R_j^{2}}, \]

with \(R_j^{2}\) the coefficient of determination from regressing \(x_j\) on all the other regressors. It is the factor by which \(\operatorname{Var}(\hat\beta_j)\) is multiplied, relative to what it would be with orthogonal regressors. \(VIF > 10\) — that is \(R_j^{2} > 0.9\) — is the usual warning line.

EXAMPLE 4.2 — VARIANCE INFLATION IN A SMALL DESIGN

Given. Three regressors observed on five units:

unit12345
\(x_1\)12345
\(x_2\)246811
\(x_3\)13254

Step 1 — look before computing. \(x_2\) is exactly \(2x_1\) at the first four units and departs only at the fifth, where it is \(11\) rather than \(10\). So the collinearity is near-exact but not exact, and \(X'X\) is non-singular — just barely.

Step 2 — the pairwise correlations.

\[ r(x_1, x_2) = 0.995893, \qquad r(x_1, x_3) = 0.800000, \qquad r(x_2, x_3) = 0.769554. \]

Step 3 — the variance inflation factors. Each comes from regressing one column on the other two:

regressor\(R_j^{2}\)\(VIF_j = 1/(1 - R_j^{2})\)reading
\(x_1\)0.99457286184.2593severe
\(x_2\)0.99385246162.6667severe
\(x_3\)0.730000003.7037acceptable

Step 4 — read what the numbers say. The variance of \(\hat\beta_1\) is \(184\) times larger than it would be with orthogonal regressors, so its standard error is \(\sqrt{184.2593} = 13.57\) times larger. A coefficient would have to be enormous to clear that standard error. Meanwhile \(x_3\), with \(VIF = 3.70\), is not part of the problem at all.

Step 5 — note what is not wrong. The estimates remain unbiased and the model's predictions remain good inside the range of the data. Multicollinearity damages the separation of effects, not the fit. That is why dropping one of \(x_1, x_2\) usually costs almost nothing in \(R^{2}\) while transforming the standard errors.

Interpretation. The remedies follow from the diagnosis: drop one of the pair, combine them into a single index, collect data that breaks the pattern, or accept a little bias for a large variance reduction by using ridge regression. What does not help is a larger sample of the same shape — the problem is in the design, not the sample size.

6. Importance and Applications of the GLM

WHERE THIS ONE MODEL REACHES
ApplicationWhat \(X\) holdsWhere on this site
Simple and multiple regressiona column of ones and the regressors Statistical Methods, Unit 4
Analysis of varianceindicator columns, rank deficient Design and Analysis of Experiments
Analysis of covarianceindicators and continuous columns together Design and Analysis of Experiments (STS-203)
Polynomial and response surface modelspowers and cross-products of the inputs Design and Analysis of Experiments (STS-203)
Weighted least squares\(V\) diagonal with unequal variances Econometrics, Unit 3
Time series with autocorrelated errors\(V\) with an AR(1) structure Econometrics, Unit 5

Each of these is the same estimator, the same theorem and the same geometry, with a different \(X\) and a different \(V\). That economy is the reason the general form is worth learning once rather than learning six special cases.

Key Take-aways