Skip to the content

Topics Covered

OLS Estimators OLS t-test for individual βⱼ F-test for overall significance Coefficient of Determination Adjusted R^2 WLS Special Case Symptoms Detection Remedies Consequences Durbin–Watson Statistic

Topic Overview — What & Why

Unit VI is the most application-rich part of the syllabus. Linear models describe how a response variable depends on predictors; econometrics adapts these models to economic data where classical assumptions often fail.

  • Simple & multiple linear regression: fits a line / hyperplane through data; OLS minimises sum of squared residuals. The first model students meet, and the workhorse of applied work.
  • Gauss-Markov theorem: under classical assumptions, OLS is BLUE — best linear unbiased estimator. Justifies using OLS over alternatives.
  • ML estimation in regression: assuming Gaussian errors, OLS coincides with MLE for $\boldsymbol\beta$, while $\hat\sigma^2_{ML}$ is biased (uses $n$, not $n-p$).
  • Tests on regression coefficients: $t$-tests for individual significance, $F$-tests for joint significance — how we decide which predictors matter.
  • $R^2$ & adjusted $R^2$: goodness-of-fit measures; adjusted version penalises over-fitting.
  • Tests of linear hypothesis: general $C\boldsymbol\beta=\mathbf d$ — embraces tests of equality of coefficients, joint zero restrictions, structural breaks (Chow).
  • GLS & WLS: when error variance is non-spherical, OLS is inefficient; GLS restores BLUE status.
  • Dummy variables: incorporate categorical predictors via 0/1 indicators — with one less dummy than the number of categories.
  • Multicollinearity: high correlation among predictors inflates standard errors; diagnosed by VIF and addressed via ridge or variable selection.
  • Heteroscedasticity: error variance depends on $X$; OLS unbiased but inefficient. Tests: Breusch-Pagan, White.
  • Autocorrelation & Durbin-Watson: serial correlation among residuals (common in time-series); biases OLS standard errors.
  • Logistic regression: for binary outcomes; models log-odds as linear in predictors. MLE used.
  • Restricted regression: incorporate prior linear restrictions (deterministic, stochastic, mixed) on $\boldsymbol\beta$ — reduces variance when restrictions hold.
  • Stochastic regressors & errors-in-variables: when $X$ is correlated with error, OLS is biased — classical "attenuation" toward zero.
  • Instrumental variables & 2SLS: remedy for endogenous regressors using exogenous instruments; basis of causal inference in econometrics.
  • Simultaneous equations & identification: in multi-equation systems, identification rules (order & rank conditions) determine whether parameters can be recovered.
  • $k$-class estimator: family that nests OLS ($k=0$) and 2SLS ($k=1$); LIML chooses $k$ as smallest root of a generalized eigenvalue problem.

1. Simple Linear Regression

Why this section? Simple linear regression is the simplest non-trivial statistical model and contains in miniature all the ideas that recur in multiple regression, GLM, and econometrics.

Model: $y_i=\beta_0+\beta_1 x_i+\varepsilon_i,\,\varepsilon_i\overset{\text{iid}}{\sim}N(0,\sigma^2),\,i=1,\ldots,n.$

OLS Estimators

$$\hat\beta_1=\frac{S_{xy}}{S_{xx}}=\frac{\sum(x_i-\bar x)(y_i-\bar y)}{\sum(x_i-\bar x)^2},\quad \hat\beta_0=\bar y-\hat\beta_1\bar x.$$

Properties

response y predictor x ŷ = b₀ + b₁x eᵢ Least-squares line minimises Σ eᵢ²
Ordinary least squares. The fitted line is chosen so that the sum of squared vertical residuals $e_i=y_i-\hat y_i$ (orange) is as small as possible.
EXAMPLE 1 Data $(1,2),(2,3),(3,5),(4,4),(5,6).$ $\bar x=3,\bar y=4,\,S_{xx}=10,S_{xy}=9.$ $\hat\beta_1=0.9,\hat\beta_0=4-2.7=1.3.$ Line: $\hat y=1.3+0.9x.$
EXAMPLE 2 $n=20,\bar x=10,S_{xx}=100,\hat\beta_1=2,\hat\sigma^2=4.$ $\text{SE}(\hat\beta_1)=\sqrt{4/100}=0.2,$ 95% CI: $2\pm 2.10\cdot 0.2=(1.58,2.42).$

🌍 Where it's used in real life

  1. Predicting sales from advertising spend.
  2. House price from floor area.
  3. Crop yield from rainfall.
  4. Exam score from study hours.
  5. Fuel use from distance driven.

2. Multiple Linear Regression

Matrix form: $\mathbf y=\mathbf X\boldsymbol\beta+\boldsymbol\varepsilon$ where $\mathbf X$ is $n\times p,\,\boldsymbol\varepsilon\sim N(\mathbf 0,\sigma^2 I).$

OLS

$$\hat{\boldsymbol\beta}=(\mathbf X^T\mathbf X)^{-1}\mathbf X^T\mathbf y,\quad \hat{\mathbf y}=\mathbf H\mathbf y,\quad \mathbf H=\mathbf X(\mathbf X^T\mathbf X)^{-1}\mathbf X^T.$$ $\mathbf H$ is the hat (projection) matrix; $\mathbf H^2=\mathbf H,\,\mathbf H^T=\mathbf H,\,\text{tr}(\mathbf H)=p.$

Properties

EXAMPLE 1 Two regressors plus intercept, $n=10$. $\hat{\boldsymbol\beta}=(X'X)^{-1}X'y.$ Compute $X'X$ as $3\times 3$ matrix and invert.
EXAMPLE 2 For a balanced design with orthogonal columns: $X'X$ diagonal $\Rightarrow$ regression coefficients independent.

🌍 Where it's used in real life

  1. House price from size, location and age.
  2. Salary from education, experience and skills.
  3. Demand from price, income and ads.
  4. Health outcome from several risk factors.
  5. Yield from rain, fertiliser and temperature.

3. Gauss–Markov Theorem

Under linear model with $E(\boldsymbol\varepsilon)=\mathbf 0$ and $\text{Var}(\boldsymbol\varepsilon)=\sigma^2 I$, the OLS estimator $\hat{\boldsymbol\beta}$ is the Best Linear Unbiased Estimator (BLUE) of $\boldsymbol\beta$ — minimum variance among all linear unbiased estimators.

Linearity in $\mathbf y$: $\hat\beta=A\mathbf y$ for $A=(X'X)^{-1}X'.$ "Best" in the sense that for any linear estimator $\tilde\beta=B\mathbf y$ that is unbiased, $\text{Var}(\tilde\beta)-\text{Var}(\hat\beta)$ is positive semi-definite.

EXAMPLE 1 Mean of iid sample: $\bar X$ is BLUE of $\mu$ — a special case where $X=\mathbf 1.$
EXAMPLE 2 For SLR, $\hat\beta_1=\sum c_i y_i$ with $c_i=(x_i-\bar x)/S_{xx}$ — minimum variance among linear unbiased.

🌍 Where it's used in real life

  1. Justifying OLS as the best linear method.
  2. Reliable slope estimates in forecasting.
  3. Engineering calibration lines.
  4. Trusting regression results in research.
  5. A baseline to compare other estimators.

4. Maximum Likelihood Estimation

Under normality $\boldsymbol\varepsilon\sim N(0,\sigma^2 I)$: $$\hat{\boldsymbol\beta}_{ML}=\hat{\boldsymbol\beta}_{OLS},\qquad \hat\sigma^2_{ML}=\frac{\text{SSE}}{n}\quad(\text{biased}).$$ Both MLEs of $\boldsymbol\beta$ and $\sigma^2$ are jointly sufficient.

EXAMPLE 1 $n=20,\text{SSE}=80,p=3.$ $\hat\sigma^2_{ML}=4,$ $\hat\sigma^2_{unbiased}=80/17\approx 4.71.$
EXAMPLE 2 Under non-normal errors, OLS is still BLUE but no longer MLE.

🌍 Where it's used in real life

  1. Fitting regression with normal errors.
  2. Foundation for logistic and other GLMs.
  3. Credit-risk scoring.
  4. Estimating error variance for prediction intervals.
  5. Maximum-likelihood curve fitting.

5. Hypothesis Tests on Regression Coefficients

$t$-test for individual $\beta_j$

$$t=\frac{\hat\beta_j-\beta_j^0}{\text{SE}(\hat\beta_j)}\sim t_{n-p}.$$

$F$-test for overall significance

$H_0:\beta_1=\cdots=\beta_{p-1}=0$: $$F=\frac{\text{SSR}/(p-1)}{\text{SSE}/(n-p)}\sim F_{p-1,n-p}.$$
EXAMPLE 1 $\hat\beta_1=2.5,\,\text{SE}=0.5,n=15,p=3.$ $t=5,$ $df=12,$ $p<0.001.$
EXAMPLE 2 $R^2=0.8,n=30,p=4.$ $F=\frac{0.8/3}{0.2/26}=34.67,$ highly significant.

🌍 Where it's used in real life

  1. Is advertising's effect statistically real?
  2. Does drug dose matter?
  3. Which predictors to keep in a model.
  4. Is the price effect significant?
  5. Testing whether a policy had an impact.

6. ANOVA for Linear Model, $R^2$, Adjusted $R^2$

SourcedfSSMS
Regression$p-1$$\text{SSR}=\sum(\hat y_i-\bar y)^2$SSR/(p-1)
Error$n-p$$\text{SSE}=\sum(y_i-\hat y_i)^2$SSE/(n-p)
Total$n-1$$\text{SST}=\sum(y_i-\bar y)^2$

Coefficient of Determination

$$R^2=\frac{\text{SSR}}{\text{SST}}=1-\frac{\text{SSE}}{\text{SST}},\quad 0\le R^2\le 1.$$

Adjusted $R^2$

$$R_{adj}^2=1-(1-R^2)\frac{n-1}{n-p}.$$ Penalizes adding non-informative regressors.
EXAMPLE 1 $n=25,p=4,R^2=0.85.$ $R_{adj}^2=1-0.15(24/21)=1-0.171=0.829.$
EXAMPLE 2 Adding irrelevant variable: $R^2$ never decreases but $R_{adj}^2$ may decrease — useful for model selection.

🌍 Where it's used in real life

  1. How much variation the model explains (R²).
  2. Overall significance of a model (F-test).
  3. Comparing the fit of competing models.
  4. Reporting predictive power to stakeholders.
  5. Comparing feature sets in machine learning.

7. Tests of Linear Hypothesis

$H_0:C\boldsymbol\beta=\mathbf d$ where $C$ is $q\times p$ of full row rank. $$F=\frac{(C\hat\beta-d)^T[C(X'X)^{-1}C^T]^{-1}(C\hat\beta-d)/q}{\text{SSE}/(n-p)}\sim F_{q,n-p}.$$ Equivalently, fit restricted (SSE_R) and full (SSE_F) models: $$F=\frac{(\text{SSE}_R-\text{SSE}_F)/q}{\text{SSE}_F/(n-p)}.$$

EXAMPLE 1 Test $H_0:\beta_2=\beta_3$ in 4-coefficient model: $C=(0,0,1,-1),\,d=0.$
EXAMPLE 2 Chow test for structural break: full model with separate slopes vs pooled. $F$-test compares restricted vs full SSE.

🌍 Where it's used in real life

  1. Testing constant returns to scale in economics.
  2. Chow test for a structural break.
  3. Testing whether two coefficients are equal.
  4. Joint significance of a group of variables.
  5. Testing the effect of a policy change.

8. Generalized & Weighted Least Squares

If $\text{Var}(\boldsymbol\varepsilon)=\sigma^2 V$ with known $V$ (PD): $$\hat{\boldsymbol\beta}_{GLS}=(X^T V^{-1}X)^{-1}X^T V^{-1}y.$$ This is BLUE under heteroscedasticity/autocorrelation.

WLS Special Case

$V=\text{diag}(1/w_i)$, weights $w_i\propto 1/\sigma_i^2$: $$\hat\beta_{WLS}=\arg\min\sum w_i(y_i-x_i^T\beta)^2.$$
EXAMPLE 1 If $\sigma_i^2=\sigma^2 x_i^2$, use $w_i=1/x_i^2$ — equivalent to regressing $y_i/x_i$ on $1/x_i$ and constant on $x_i.$
EXAMPLE 2 With AR(1) errors, $V$ has Toeplitz form; GLS gives Cochrane–Orcutt-like transformation.

🌍 Where it's used in real life

  1. Data whose error spread is unequal.
  2. Time series with correlated errors.
  3. Weighted survey data.
  4. Grouped/panel data with heteroscedasticity.
  5. Volatility-adjusted financial regression.

9. Indicator/Dummy Variables

Encode categorical predictors with $k$ levels using $k-1$ binary dummies (avoid dummy variable trap from perfect multicollinearity with intercept).

Interpretation

Coefficient of dummy $D_j$ measures effect relative to baseline category.
EXAMPLE 1 Wage $=\beta_0+\beta_1\text{Edu}+\beta_2\text{Female}.$ $\beta_2$ is wage gap between female (=1) and male (=0) at same education.
EXAMPLE 2 Seasonal dummies: 4 quarters, use $D_2,D_3,D_4$ with Q1 as base. Test $D_2=D_3=D_4=0$ for absence of seasonality.

🌍 Where it's used in real life

  1. Adding gender or region to a wage model.
  2. Seasonal (quarterly) effects in sales.
  3. A before/after policy indicator.
  4. Product-category effects on price.
  5. Treatment vs control group indicator.

10. Multicollinearity

High linear dependence among regressors → $X'X$ near-singular.

Symptoms

Intuition. When two predictors move together, the data cannot tell whose effect is whose, so the individual slopes are estimated imprecisely (huge standard errors) even though the predictors jointly explain the response well (high $R^2$, significant $F$). This is the classic “insignificant $t$'s but significant $F$” signature — the estimates are unbiased but unstable, not wrong on average.

Detection

Remedies

EXAMPLE 1 $R_j^2=0.95\Rightarrow\text{VIF}=20$ — severe.
EXAMPLE 2 Including both $\text{Income}$ and $\text{Income}^2$: high correlation but justified — center variable to reduce VIF.

🌍 Where it's used in real life

  1. Untangling correlated predictors (height and weight).
  2. Economic indicators that move together.
  3. Overlapping marketing channels.
  4. Sensors that duplicate information.
  5. Deciding which correlated feature to drop.

11. Heteroscedasticity

$\text{Var}(\varepsilon_i)=\sigma_i^2$ varies across $i.$

Consequences

Tests

Remedies

WLS, robust (White) standard errors: $\hat V(\hat\beta)=(X'X)^{-1}X'\hat\Omega X(X'X)^{-1},\,\hat\Omega=\text{diag}(\hat e_i^2).$
EXAMPLE 1 Cross-section of households: variance of expenditure rises with income → heteroscedastic.
EXAMPLE 2 Goldfeld–Quandt: $n=60$, sort by $X$, drop middle 12, fit two regressions of size 24. $F$-stat $=S_{high}^2/S_{low}^2.$

🌍 Where it's used in real life

  1. Spending that varies more as income rises.
  2. Profit variability that grows with firm size.
  3. Correcting standard errors in economics.
  4. Financial data with changing volatility.
  5. Robust inference in cross-section studies.

12. Autocorrelation & Durbin–Watson

$\text{Cov}(\varepsilon_t,\varepsilon_s)\ne 0$ for $t\ne s$ — common in time series. AR(1): $\varepsilon_t=\rho\varepsilon_{t-1}+u_t.$

Consequences

OLS unbiased but inefficient; SEs biased.

Durbin–Watson Statistic

$$d=\frac{\sum_{t=2}^n(\hat e_t-\hat e_{t-1})^2}{\sum_{t=1}^n\hat e_t^2}\approx 2(1-\hat\rho).$$ $d\approx 2$: no autocorrelation. $d<2$: positive autocorrelation. $d>2$: negative.

Cochrane–Orcutt Procedure

Estimate $\hat\rho$ from residuals; transform variables: $y_t-\hat\rho y_{t-1}=\beta_0(1-\hat\rho)+\beta_1(x_t-\hat\rho x_{t-1})+u_t.$ Iterate.
EXAMPLE 1 $\hat\rho=0.5\Rightarrow d\approx 1.0$ — significant positive autocorrelation.
EXAMPLE 2 $d=2.4$ with $n=30,k=3,$ $d_L=1.21,d_U=1.65$ — no autocorrelation (since $4-d=1.6$ between $d_L$ and $d_U$ — inconclusive on negative side).

🌍 Where it's used in real life

  1. Correlated errors in monthly sales.
  2. Stock-return regressions.
  3. GDP and inflation time series.
  4. Weather-driven demand models.
  5. Spotting a missing trend or lag variable.

13. Logistic Regression

For binary $y\in\{0,1\}$: $P(y=1\mid x)=\pi(x)=\frac{e^{x'\beta}}{1+e^{x'\beta}}.$ Logit: $\log\frac{\pi}{1-\pi}=x'\beta.$

Estimation

By maximum likelihood (no closed form). Newton–Raphson / IRLS.

Interpretation

$e^{\beta_j}$ = odds ratio for unit increase in $x_j.$

Wald Test

$z=\hat\beta_j/\text{SE}(\hat\beta_j)\sim N(0,1)$ asymptotically.
EXAMPLE 1 Disease vs age: $\hat\beta_{age}=0.05.$ Odds increase by $e^{0.05}\approx 5.13\%$ per year.
EXAMPLE 2 Two predictors: $\text{logit}(\pi)=-2+0.05\,\text{age}+0.4\,\text{smoke}.$ At age 50, smoker: $\text{logit}=-2+2.5+0.4=0.9,\,\pi=e^{0.9}/(1+e^{0.9})=0.711.$

🌍 Where it's used in real life

  1. Predicting loan default (yes/no).
  2. Diagnosing disease present/absent.
  3. Predicting customer churn.
  4. Classifying email as spam.
  5. Click vs no-click on online ads.

14. Restricted Regression Estimation

Exact Restriction

$R\beta=r$ (exact, deterministic). Restricted estimator: $$\hat\beta_R=\hat\beta+(X'X)^{-1}R'[R(X'X)^{-1}R']^{-1}(r-R\hat\beta).$$ $\text{Var}(\hat\beta_R)\le\text{Var}(\hat\beta)$ in PSD ordering — variance reduction.

Stochastic Restriction

$r=R\beta+v,\,v\sim N(0,\Phi)$ — Theil–Goldberger Mixed Estimator: $$\hat\beta_M=(X'X+R'\Phi^{-1}R)^{-1}(X'y+R'\Phi^{-1}r).$$

Mixed Restrictions

Combine deterministic and stochastic constraints — generalized form using GLS framework.
EXAMPLE 1 Cobb-Douglas $\log Y=\beta_0+\beta_1\log K+\beta_2\log L+\varepsilon.$ Restriction $\beta_1+\beta_2=1$ (constant returns to scale).
EXAMPLE 2 Prior knowledge: $\beta_1\approx 0.5\pm 0.1.$ Use stochastic restriction with $r=0.5,R=(0,1,0),\Phi=0.01.$

🌍 Where it's used in real life

  1. Imposing economic constraints like returns to scale.
  2. Budget shares that must sum to one.
  3. Using prior knowledge to stabilise estimates.
  4. Enforcing a known sign or size on a coefficient.
  5. Blending survey data with expert priors.

15. Stochastic Regressors & Errors-in-Variables (EIV)

Stochastic Regressors

If $X$ random but independent of $\varepsilon$: OLS still unbiased and consistent. If $\text{Cov}(X,\varepsilon)\ne 0$: OLS biased and inconsistent.

Errors-in-Variables Model

Observe $X^*=X+u,\,Y=\beta_0+\beta_1 X+\varepsilon$ with $X$ unobserved. OLS based on $X^*$ gives $\text{plim}\,\hat\beta_1=\beta_1\frac{\sigma_X^2}{\sigma_X^2+\sigma_u^2}<\beta_1$ — attenuation bias toward 0.
EXAMPLE 1 Measurement error in $X$ leads to underestimated slope. With $\sigma_X^2=1,\sigma_u^2=0.5$: bias factor = $1/1.5\approx 0.67.$
EXAMPLE 2 Reverse causality (omitted variables, simultaneity) → $\text{Cov}(X,\varepsilon)\ne 0$ → IV estimation needed.

🌍 Where it's used in real life

  1. Measurement error in self-reported income.
  2. Test-score models with noisy inputs.
  3. Sensor error in engineering models.
  4. Bias from self-reported survey data.
  5. Correcting attenuation bias in econometrics.

16. Instrumental Variable (IV) Estimator

Instrument $Z$ satisfies: (i) $\text{Cov}(Z,X)\ne 0$ (relevance), (ii) $\text{Cov}(Z,\varepsilon)=0$ (exogeneity). $$\hat\beta_{IV}=(Z'X)^{-1}Z'y.$$ Consistent for $\beta$ even when $X$ is endogenous; less efficient than OLS when OLS is valid.

Intuition. When $X$ is correlated with the error, OLS cannot tell the effect of $X$ apart from the error moving with it, so it is biased. A valid instrument $Z$ shifts $X$ (relevance) only through channels unrelated to the error (exogeneity); using just that “clean” variation in $X$ recovers a consistent slope. Weak instruments (low $\text{Cov}(Z,X)$) make this cure unreliable.

EXAMPLE 1 Returns to schooling: education $X$ correlated with ability ($\varepsilon$). Use distance-to-college as instrument $Z.$
EXAMPLE 2 Demand–supply: log price endogenous in demand. Use weather (supply shifter) as instrument.

🌍 Where it's used in real life

  1. Returns to schooling using distance to college.
  2. Demand estimation using weather as a supply shifter.
  3. Policy effects using eligibility rules.
  4. Health effect of a habit using taxes as an instrument.
  5. Causal effects when a regressor is endogenous.

17. Simultaneous Equations Model & Identification

Multiple endogenous variables determined jointly: $$Y\Gamma+XB=U,$$ $Y$: $T\times M$ endogenous, $X$: $T\times K$ exogenous.

Reduced Form

$Y=X\Pi+V,\,\Pi=-B\Gamma^{-1}.$ Always estimable by OLS.

Identification (Order Condition)

For an equation with $m$ included endogenous and $k$ included exogenous variables:

Rank Condition

The matrix of excluded variables' coefficients in other equations has rank $M-1$ — necessary and sufficient.
EXAMPLE 1 Demand: $Q=\alpha_0+\alpha_1 P+\alpha_2 Y+u_1$; Supply: $Q=\beta_0+\beta_1 P+\beta_2 W+u_2.$ Each equation excludes one exogenous variable; both just-identified.
EXAMPLE 2 If supply equation has no excluded exogenous variable, demand cannot be identified by IV — order condition fails.

🌍 Where it's used in real life

  1. Supply and demand set together.
  2. Macro models of income and consumption.
  3. Wage–price systems.
  4. Market-equilibrium modelling.
  5. Interdependent policy variables.

18. 2SLS & k-class Estimator

Two-Stage Least Squares (2SLS)

  1. Regress each endogenous regressor on all exogenous variables → get fitted values $\hat X.$
  2. Replace endogenous regressor with $\hat X$ in original equation, run OLS.
Equivalent to IV with $Z$ being all exogenous variables in the system.

k-class Estimator

Generalized form: $$\hat\beta_k=(X'(I-kM)X)^{-1}X'(I-kM)y,$$ where $M=I-Z(Z'Z)^{-1}Z'.$
EXAMPLE 1 Demand $Q=\alpha+\beta P+\gamma Y+u.$ With instruments (income, weather), regress $P$ on these → $\hat P$, then OLS of $Q$ on $\hat P,Y.$
EXAMPLE 2 LIML preferred over 2SLS in small samples and weak instrument settings — chooses $k$ as smallest root of generalized eigenvalue equation.

🌍 Where it's used in real life

  1. Estimating demand with an endogenous price.
  2. Labour-supply and wage systems.
  3. Estimating macro-econometric models.
  4. Causal inference in economics.
  5. Robust estimation with weak instruments (LIML).