Skip to the content

Topics Covered

Disturbance Assumptions AR(1) Scheme Consequences Durbin–Watson Breusch–Godfrey Estimating \(\rho\) Cochrane–Orcutt Newey–West SE
On this page
  1. 1. Disturbance Term and Its Assumptions
  2. 2. Consequences of Autocorrelated Disturbances
  3. 3. Detecting Autocorrelation
  4. 4. Estimating the Autocorrelation Coefficient \(\rho\)
  5. 5. Solutions for Autocorrelation
  6. 6. When the Model Contains a Lagged Dependent Variable
  7. 7. Simultaneous-Equation Models
  8. Key Take-aways from Unit 5

1. Disturbance Term and Its Assumptions

CLASSICAL ASSUMPTION ON \(u\) \[ \mathrm{Cov}(u_t, u_s) = 0 \quad \text{for all } t \ne s. \]

This means disturbances are serially uncorrelated — what happened in period \(t\) tells us nothing about what will happen in period \(s\).

1.1 What Is Autocorrelation?

Autocorrelation (serial correlation) means the assumption above is violated: \(\mathrm{Cov}(u_t, u_s) \ne 0\) for some \(t \ne s\). It is the time-series analogue of heteroscedasticity.

1.2 The First-Order Autoregressive (AR(1)) Scheme

\[ u_t = \rho u_{t-1} + \varepsilon_t, \qquad |\rho| < 1,\quad \varepsilon_t \stackrel{iid}{\sim} (0, \sigma_\varepsilon^2). \]

Here \(\rho\) is the autocorrelation coefficient measuring the dependence of \(u_t\) on its own past. \(\rho > 0\) is "positive" autocorrelation (errors persist in the same direction), \(\rho < 0\) is "negative" (errors alternate sign).

1.3 Why Does Autocorrelation Arise?

2. Consequences of Autocorrelated Disturbances

  1. OLS estimators remain unbiased and consistent (provided regressors are non-stochastic or strictly exogenous).
  2. OLS is no longer efficient (not BLUE).
  3. The usual variance formula \(\sigma^2(\mathbf X'\mathbf X)^{-1}\) is wrong; standard errors are typically underestimated, making \(t\)-statistics misleadingly large.
  4. \(R^2\) is overstated.
  5. Hypothesis tests are unreliable.
  6. The forecast standard errors are biased.

3. Detecting Autocorrelation

3.1 Graphical Method

Plot residuals \(\hat u_t\) against time, or plot \(\hat u_t\) vs. \(\hat u_{t-1}\). A clear upward or downward pattern across time, or a clear positive (negative) slope in the lag plot, indicates positive (negative) autocorrelation.

No autocorrelation (random) Positive autocorrelation (runs) time t → time t → residual
Fig 5.1 — Residuals plotted against time. When they scatter randomly around zero, switching sign often (left), there is no autocorrelation. When they move in smooth runs — several positives followed by several negatives (right) — successive errors are correlated: positive autocorrelation, which the Durbin–Watson test detects (\(d \approx 2(1-\hat\rho)\), so \(d\) well below 2).

3.2 The Durbin–Watson Test

DURBIN–WATSON STATISTIC \[ d = \frac{\sum_{t=2}^{n} (\hat u_t - \hat u_{t-1})^2}{\sum_{t=1}^{n} \hat u_t^2}. \]

Approximate relation with the sample autocorrelation coefficient \(\hat\rho\):

\[ d \;\approx\; 2(1 - \hat\rho). \]

So:

3.3 Decision Rules

Durbin and Watson provide critical values \(d_L\) (lower) and \(d_U\) (upper) depending on \(n\) and \(k\). The decision zones are:

Range of \(d\)Conclusion
\(0 < d < d_L\)Reject \(H_0\) — positive autocorrelation.
\(d_L \le d \le d_U\)Inconclusive.
\(d_U < d < 4 - d_U\)Do not reject — no autocorrelation.
\(4 - d_U \le d \le 4 - d_L\)Inconclusive.
\(4 - d_L < d < 4\)Reject \(H_0\) — negative autocorrelation.

3.4 Assumptions of the DW Test

3.5 Breusch–Godfrey (LM) Test

More general than DW — tests for autocorrelation of any order \(p\), allows lagged dependent variables. Procedure:

  1. Run OLS on the original model; get residuals \(\hat u_t\).
  2. Regress \(\hat u_t\) on \(\mathbf X_t\) and \(\hat u_{t-1}, \ldots, \hat u_{t-p}\); get \(R^2\) of this auxiliary regression.
  3. Statistic: \((n-p) R^2 \sim \chi^2_p\) under \(H_0\) of no autocorrelation.

4. Estimating the Autocorrelation Coefficient \(\rho\)

4.1 From the Durbin–Watson Statistic

\[ \hat\rho \;\approx\; 1 - d/2. \]

Fast back-of-envelope estimate.

4.2 From Residual Regression

\[ \hat\rho = \frac{\sum_{t=2}^{n} \hat u_t \hat u_{t-1}}{\sum_{t=2}^{n} \hat u_{t-1}^2}. \]

The OLS slope from regressing \(\hat u_t\) on \(\hat u_{t-1}\) (no intercept).

4.3 Cochrane–Orcutt Iterative Procedure

  1. Run OLS on the original model; get \(\hat u_t\).
  2. Estimate \(\hat\rho\) from the residual regression above.
  3. Transform the variables: \(Y_t^* = Y_t - \hat\rho Y_{t-1}\), \(X_t^* = X_t - \hat\rho X_{t-1}\).
  4. Apply OLS to the transformed model.
  5. Repeat with the new residuals until \(\hat\rho\) converges.

This is a feasible Generalised Least Squares (FGLS) estimator under the AR(1) assumption.

5. Solutions for Autocorrelation

5.1 Generalised Least Squares (GLS)

When \(\mathrm{Var}(\mathbf u) = \sigma^2 \boldsymbol\Omega\) for known \(\boldsymbol\Omega\):

\[ \hat{\boldsymbol\beta}_{\text{GLS}} = (\mathbf X' \boldsymbol\Omega^{-1} \mathbf X)^{-1}\, \mathbf X' \boldsymbol\Omega^{-1} \mathbf Y. \]

For AR(1) errors, \(\boldsymbol\Omega\) has the Prais–Winsten / Cochrane–Orcutt structure.

5.2 Newey–West (HAC) Standard Errors

Heteroscedasticity- and Autocorrelation-Consistent (HAC) standard errors leave OLS point estimates unchanged but adjust the SE formula to be valid in the presence of arbitrary (mild) autocorrelation and heteroscedasticity. Default in modern time-series regression software.

5.3 First-Difference Transformation

If \(\rho \approx 1\) (near unit root), the simplest fix is to estimate the model in first differences:

\[ \Delta Y_t = \beta_1 \Delta X_t + \Delta u_t, \]

where \(\Delta u_t = \varepsilon_t\) is now serially uncorrelated.

5.4 Adding Lagged Dependent / Independent Variables

Sometimes autocorrelation reflects a missing dynamic — adding \(Y_{t-1}\) or \(X_{t-1}\) to the regression captures the persistence and removes the autocorrelation.

EXAMPLE 1 — Durbin–Watson interpretation

A consumption regression on \(n = 25\) annual observations with \(k = 2\) regressors gives \(d = 0.85\). Tables: \(d_L = 1.21, d_U = 1.55\) at 5%. Since \(0.85 < 1.21\), reject \(H_0\) — strong positive autocorrelation. Implied \(\hat\rho \approx 1 - 0.85/2 = 0.575\).

EXAMPLE 2 — Cochrane–Orcutt step-by-step

Initial regression: \(Y_t = \hat\beta_0 + \hat\beta_1 X_t + \hat u_t\); \(\hat\rho = 0.60\) from residuals.

Transform: \(Y_t^* = Y_t - 0.60 Y_{t-1}\), \(X_t^* = X_t - 0.60 X_{t-1}\), \((1 - 0.60) = 0.40\) as the new intercept multiplier.

Re-run OLS on transformed data. Compute new residuals \(\hat u_t^{(1)}\), get new \(\hat\rho^{(1)}\). Iterate. The procedure converges in 2–4 iterations for most well-behaved models. The final estimates have valid standard errors and the DW statistic of the transformed model should be near 2.

6. When the Model Contains a Lagged Dependent Variable

The standard Durbin–Watson test is biased toward 2 when the regression includes a lagged dependent variable like \(Y_{t-1}\). In that case use Durbin's h-test:

\[ h = \hat\rho \sqrt{\frac{n}{1 - n \cdot \hat{\mathrm{Var}}(\hat\beta_{Y_{t-1}})}}, \]

which is asymptotically standard normal under \(H_0: \rho = 0\). Compare \(|h|\) with 1.96 for a two-tailed 5% test.

7. Simultaneous-Equation Models

In many economic systems variables are jointly determined — e.g. price and quantity are fixed together by demand and supply. Such a system is a set of structural equations; a variable that is explained within the system is endogenous, one determined outside it is exogenous.

\[ \text{Demand: } Q = \alpha_0 + \alpha_1 P + \alpha_2 I + u_1, \qquad \text{Supply: } Q = \beta_0 + \beta_1 P + u_2. \]

Simultaneity bias: because \(P\) is correlated with the disturbances, applying OLS to a structural equation gives biased and inconsistent estimates. Solving the system for the endogenous variables in terms of exogenous ones gives the reduced form, whose OLS estimates are consistent.

THE IDENTIFICATION PROBLEM

Whether the structural parameters can be recovered from the reduced form is the identification problem. The order condition (necessary): an equation is identified if the number of exogenous variables excluded from it is at least \((\text{number of endogenous variables}) - 1\), i.e. \(K - k \ge m - 1\).

2SLS replaces the endogenous regressor by its fitted value from a first-stage regression on all exogenous variables (instruments), then applies OLS — removing the correlation with the disturbance. In the demand–supply example, income \(I\) (excluded from supply) identifies the supply equation.

Key Take-aways from Unit 5