This means disturbances are serially uncorrelated — what happened in period \(t\) tells us nothing about what will happen in period \(s\).
Autocorrelation (serial correlation) means the assumption above is violated: \(\mathrm{Cov}(u_t, u_s) \ne 0\) for some \(t \ne s\). It is the time-series analogue of heteroscedasticity.
Here \(\rho\) is the autocorrelation coefficient measuring the dependence of \(u_t\) on its own past. \(\rho > 0\) is "positive" autocorrelation (errors persist in the same direction), \(\rho < 0\) is "negative" (errors alternate sign).
Plot residuals \(\hat u_t\) against time, or plot \(\hat u_t\) vs. \(\hat u_{t-1}\). A clear upward or downward pattern across time, or a clear positive (negative) slope in the lag plot, indicates positive (negative) autocorrelation.
Approximate relation with the sample autocorrelation coefficient \(\hat\rho\):
\[ d \;\approx\; 2(1 - \hat\rho). \]So:
Durbin and Watson provide critical values \(d_L\) (lower) and \(d_U\) (upper) depending on \(n\) and \(k\). The decision zones are:
| Range of \(d\) | Conclusion |
|---|---|
| \(0 < d < d_L\) | Reject \(H_0\) — positive autocorrelation. |
| \(d_L \le d \le d_U\) | Inconclusive. |
| \(d_U < d < 4 - d_U\) | Do not reject — no autocorrelation. |
| \(4 - d_U \le d \le 4 - d_L\) | Inconclusive. |
| \(4 - d_L < d < 4\) | Reject \(H_0\) — negative autocorrelation. |
More general than DW — tests for autocorrelation of any order \(p\), allows lagged dependent variables. Procedure:
Fast back-of-envelope estimate.
The OLS slope from regressing \(\hat u_t\) on \(\hat u_{t-1}\) (no intercept).
This is a feasible Generalised Least Squares (FGLS) estimator under the AR(1) assumption.
When \(\mathrm{Var}(\mathbf u) = \sigma^2 \boldsymbol\Omega\) for known \(\boldsymbol\Omega\):
\[ \hat{\boldsymbol\beta}_{\text{GLS}} = (\mathbf X' \boldsymbol\Omega^{-1} \mathbf X)^{-1}\, \mathbf X' \boldsymbol\Omega^{-1} \mathbf Y. \]For AR(1) errors, \(\boldsymbol\Omega\) has the Prais–Winsten / Cochrane–Orcutt structure.
Heteroscedasticity- and Autocorrelation-Consistent (HAC) standard errors leave OLS point estimates unchanged but adjust the SE formula to be valid in the presence of arbitrary (mild) autocorrelation and heteroscedasticity. Default in modern time-series regression software.
If \(\rho \approx 1\) (near unit root), the simplest fix is to estimate the model in first differences:
where \(\Delta u_t = \varepsilon_t\) is now serially uncorrelated.
Sometimes autocorrelation reflects a missing dynamic — adding \(Y_{t-1}\) or \(X_{t-1}\) to the regression captures the persistence and removes the autocorrelation.
A consumption regression on \(n = 25\) annual observations with \(k = 2\) regressors gives \(d = 0.85\). Tables: \(d_L = 1.21, d_U = 1.55\) at 5%. Since \(0.85 < 1.21\), reject \(H_0\) — strong positive autocorrelation. Implied \(\hat\rho \approx 1 - 0.85/2 = 0.575\).
Initial regression: \(Y_t = \hat\beta_0 + \hat\beta_1 X_t + \hat u_t\); \(\hat\rho = 0.60\) from residuals.
Transform: \(Y_t^* = Y_t - 0.60 Y_{t-1}\), \(X_t^* = X_t - 0.60 X_{t-1}\), \((1 - 0.60) = 0.40\) as the new intercept multiplier.
Re-run OLS on transformed data. Compute new residuals \(\hat u_t^{(1)}\), get new \(\hat\rho^{(1)}\). Iterate. The procedure converges in 2–4 iterations for most well-behaved models. The final estimates have valid standard errors and the DW statistic of the transformed model should be near 2.
The standard Durbin–Watson test is biased toward 2 when the regression includes a lagged dependent variable like \(Y_{t-1}\). In that case use Durbin's h-test:
which is asymptotically standard normal under \(H_0: \rho = 0\). Compare \(|h|\) with 1.96 for a two-tailed 5% test.
In many economic systems variables are jointly determined — e.g. price and quantity are fixed together by demand and supply. Such a system is a set of structural equations; a variable that is explained within the system is endogenous, one determined outside it is exogenous.
Simultaneity bias: because \(P\) is correlated with the disturbances, applying OLS to a structural equation gives biased and inconsistent estimates. Solving the system for the endogenous variables in terms of exogenous ones gives the reduced form, whose OLS estimates are consistent.
Whether the structural parameters can be recovered from the reduced form is the identification problem. The order condition (necessary): an equation is identified if the number of exogenous variables excluded from it is at least \((\text{number of endogenous variables}) - 1\), i.e. \(K - k \ge m - 1\).
2SLS replaces the endogenous regressor by its fitted value from a first-stage regression on all exogenous variables (instruments), then applies OLS — removing the correlation with the disturbance. In the demand–supply example, income \(I\) (excluded from supply) identifies the supply equation.