Topics Covered
Contents
- 1. Simple Linear Regression
- 2. Multiple Linear Regression
- 3. Gauss–Markov Theorem
- 4. ML Estimation in Regression
- 5. Hypothesis Tests on Regression Coefficients
- 6. ANOVA for Linear Model, $R^2$, Adjusted $R^2$
- 7. Tests of Linear Hypothesis
- 8. GLS & Weighted Least Squares
- 9. Indicator/Dummy Variables
- 10. Multicollinearity
- 11. Heteroscedasticity
- 12. Autocorrelation & Durbin–Watson
- 13. Logistic Regression
- 14. Restricted Regression
- 15. Stochastic Regressors & EIV
- 16. Instrumental Variable Estimator
- 17. Simultaneous Equations & Identification
- 18. 2SLS & k-class Estimator
Topic Overview — What & Why
Unit VI is the most application-rich part of the syllabus. Linear models describe how a response variable depends on predictors; econometrics adapts these models to economic data where classical assumptions often fail.
- Simple & multiple linear regression: fits a line / hyperplane through data; OLS minimises sum of squared residuals. The first model students meet, and the workhorse of applied work.
- Gauss-Markov theorem: under classical assumptions, OLS is BLUE — best linear unbiased estimator. Justifies using OLS over alternatives.
- ML estimation in regression: assuming Gaussian errors, OLS coincides with MLE for $\boldsymbol\beta$, while $\hat\sigma^2_{ML}$ is biased (uses $n$, not $n-p$).
- Tests on regression coefficients: $t$-tests for individual significance, $F$-tests for joint significance — how we decide which predictors matter.
- $R^2$ & adjusted $R^2$: goodness-of-fit measures; adjusted version penalises over-fitting.
- Tests of linear hypothesis: general $C\boldsymbol\beta=\mathbf d$ — embraces tests of equality of coefficients, joint zero restrictions, structural breaks (Chow).
- GLS & WLS: when error variance is non-spherical, OLS is inefficient; GLS restores BLUE status.
- Dummy variables: incorporate categorical predictors via 0/1 indicators — with one less dummy than the number of categories.
- Multicollinearity: high correlation among predictors inflates standard errors; diagnosed by VIF and addressed via ridge or variable selection.
- Heteroscedasticity: error variance depends on $X$; OLS unbiased but inefficient. Tests: Breusch-Pagan, White.
- Autocorrelation & Durbin-Watson: serial correlation among residuals (common in time-series); biases OLS standard errors.
- Logistic regression: for binary outcomes; models log-odds as linear in predictors. MLE used.
- Restricted regression: incorporate prior linear restrictions (deterministic, stochastic, mixed) on $\boldsymbol\beta$ — reduces variance when restrictions hold.
- Stochastic regressors & errors-in-variables: when $X$ is correlated with error, OLS is biased — classical "attenuation" toward zero.
- Instrumental variables & 2SLS: remedy for endogenous regressors using exogenous instruments; basis of causal inference in econometrics.
- Simultaneous equations & identification: in multi-equation systems, identification rules (order & rank conditions) determine whether parameters can be recovered.
- $k$-class estimator: family that nests OLS ($k=0$) and 2SLS ($k=1$); LIML chooses $k$ as smallest root of a generalized eigenvalue problem.
1. Simple Linear Regression
Why this section? Simple linear regression is the simplest non-trivial statistical model and contains in miniature all the ideas that recur in multiple regression, GLM, and econometrics.
OLS Estimators
$$\hat\beta_1=\frac{S_{xy}}{S_{xx}}=\frac{\sum(x_i-\bar x)(y_i-\bar y)}{\sum(x_i-\bar x)^2},\quad \hat\beta_0=\bar y-\hat\beta_1\bar x.$$Properties
- $E(\hat\beta_1)=\beta_1,\,V(\hat\beta_1)=\sigma^2/S_{xx}.$
- $E(\hat\beta_0)=\beta_0,\,V(\hat\beta_0)=\sigma^2(1/n+\bar x^2/S_{xx}).$
- $\text{Cov}(\hat\beta_0,\hat\beta_1)=-\sigma^2\bar x/S_{xx}.$
- $\hat\sigma^2=\frac{\text{SSE}}{n-2}$ unbiased for $\sigma^2.$
🌍 Where it's used in real life
- Predicting sales from advertising spend.
- House price from floor area.
- Crop yield from rainfall.
- Exam score from study hours.
- Fuel use from distance driven.
2. Multiple Linear Regression
Matrix form: $\mathbf y=\mathbf X\boldsymbol\beta+\boldsymbol\varepsilon$ where $\mathbf X$ is $n\times p,\,\boldsymbol\varepsilon\sim N(\mathbf 0,\sigma^2 I).$
OLS
$$\hat{\boldsymbol\beta}=(\mathbf X^T\mathbf X)^{-1}\mathbf X^T\mathbf y,\quad \hat{\mathbf y}=\mathbf H\mathbf y,\quad \mathbf H=\mathbf X(\mathbf X^T\mathbf X)^{-1}\mathbf X^T.$$ $\mathbf H$ is the hat (projection) matrix; $\mathbf H^2=\mathbf H,\,\mathbf H^T=\mathbf H,\,\text{tr}(\mathbf H)=p.$Properties
- $E(\hat{\boldsymbol\beta})=\boldsymbol\beta,\,\text{Var}(\hat{\boldsymbol\beta})=\sigma^2(\mathbf X^T\mathbf X)^{-1}.$
- Residuals $\mathbf e=(\mathbf I-\mathbf H)\mathbf y;\,\text{SSE}=\mathbf e^T\mathbf e.$
- $\hat\sigma^2=\text{SSE}/(n-p)$ unbiased.
🌍 Where it's used in real life
- House price from size, location and age.
- Salary from education, experience and skills.
- Demand from price, income and ads.
- Health outcome from several risk factors.
- Yield from rain, fertiliser and temperature.
3. Gauss–Markov Theorem
Linearity in $\mathbf y$: $\hat\beta=A\mathbf y$ for $A=(X'X)^{-1}X'.$ "Best" in the sense that for any linear estimator $\tilde\beta=B\mathbf y$ that is unbiased, $\text{Var}(\tilde\beta)-\text{Var}(\hat\beta)$ is positive semi-definite.
🌍 Where it's used in real life
- Justifying OLS as the best linear method.
- Reliable slope estimates in forecasting.
- Engineering calibration lines.
- Trusting regression results in research.
- A baseline to compare other estimators.
4. Maximum Likelihood Estimation
Under normality $\boldsymbol\varepsilon\sim N(0,\sigma^2 I)$: $$\hat{\boldsymbol\beta}_{ML}=\hat{\boldsymbol\beta}_{OLS},\qquad \hat\sigma^2_{ML}=\frac{\text{SSE}}{n}\quad(\text{biased}).$$ Both MLEs of $\boldsymbol\beta$ and $\sigma^2$ are jointly sufficient.
🌍 Where it's used in real life
- Fitting regression with normal errors.
- Foundation for logistic and other GLMs.
- Credit-risk scoring.
- Estimating error variance for prediction intervals.
- Maximum-likelihood curve fitting.
5. Hypothesis Tests on Regression Coefficients
$t$-test for individual $\beta_j$
$$t=\frac{\hat\beta_j-\beta_j^0}{\text{SE}(\hat\beta_j)}\sim t_{n-p}.$$$F$-test for overall significance
$H_0:\beta_1=\cdots=\beta_{p-1}=0$: $$F=\frac{\text{SSR}/(p-1)}{\text{SSE}/(n-p)}\sim F_{p-1,n-p}.$$🌍 Where it's used in real life
- Is advertising's effect statistically real?
- Does drug dose matter?
- Which predictors to keep in a model.
- Is the price effect significant?
- Testing whether a policy had an impact.
6. ANOVA for Linear Model, $R^2$, Adjusted $R^2$
| Source | df | SS | MS |
|---|---|---|---|
| Regression | $p-1$ | $\text{SSR}=\sum(\hat y_i-\bar y)^2$ | SSR/(p-1) |
| Error | $n-p$ | $\text{SSE}=\sum(y_i-\hat y_i)^2$ | SSE/(n-p) |
| Total | $n-1$ | $\text{SST}=\sum(y_i-\bar y)^2$ |
Coefficient of Determination
$$R^2=\frac{\text{SSR}}{\text{SST}}=1-\frac{\text{SSE}}{\text{SST}},\quad 0\le R^2\le 1.$$Adjusted $R^2$
$$R_{adj}^2=1-(1-R^2)\frac{n-1}{n-p}.$$ Penalizes adding non-informative regressors.🌍 Where it's used in real life
- How much variation the model explains (R²).
- Overall significance of a model (F-test).
- Comparing the fit of competing models.
- Reporting predictive power to stakeholders.
- Comparing feature sets in machine learning.
7. Tests of Linear Hypothesis
$H_0:C\boldsymbol\beta=\mathbf d$ where $C$ is $q\times p$ of full row rank. $$F=\frac{(C\hat\beta-d)^T[C(X'X)^{-1}C^T]^{-1}(C\hat\beta-d)/q}{\text{SSE}/(n-p)}\sim F_{q,n-p}.$$ Equivalently, fit restricted (SSE_R) and full (SSE_F) models: $$F=\frac{(\text{SSE}_R-\text{SSE}_F)/q}{\text{SSE}_F/(n-p)}.$$
🌍 Where it's used in real life
- Testing constant returns to scale in economics.
- Chow test for a structural break.
- Testing whether two coefficients are equal.
- Joint significance of a group of variables.
- Testing the effect of a policy change.
8. Generalized & Weighted Least Squares
If $\text{Var}(\boldsymbol\varepsilon)=\sigma^2 V$ with known $V$ (PD): $$\hat{\boldsymbol\beta}_{GLS}=(X^T V^{-1}X)^{-1}X^T V^{-1}y.$$ This is BLUE under heteroscedasticity/autocorrelation.
WLS Special Case
$V=\text{diag}(1/w_i)$, weights $w_i\propto 1/\sigma_i^2$: $$\hat\beta_{WLS}=\arg\min\sum w_i(y_i-x_i^T\beta)^2.$$🌍 Where it's used in real life
- Data whose error spread is unequal.
- Time series with correlated errors.
- Weighted survey data.
- Grouped/panel data with heteroscedasticity.
- Volatility-adjusted financial regression.
9. Indicator/Dummy Variables
Encode categorical predictors with $k$ levels using $k-1$ binary dummies (avoid dummy variable trap from perfect multicollinearity with intercept).
Interpretation
Coefficient of dummy $D_j$ measures effect relative to baseline category.🌍 Where it's used in real life
- Adding gender or region to a wage model.
- Seasonal (quarterly) effects in sales.
- A before/after policy indicator.
- Product-category effects on price.
- Treatment vs control group indicator.
10. Multicollinearity
High linear dependence among regressors → $X'X$ near-singular.
Symptoms
- Large standard errors despite high $R^2.$
- Coefficients sensitive to small data changes.
- "Wrong" sign or large magnitude.
Intuition. When two predictors move together, the data cannot tell whose effect is whose, so the individual slopes are estimated imprecisely (huge standard errors) even though the predictors jointly explain the response well (high $R^2$, significant $F$). This is the classic “insignificant $t$'s but significant $F$” signature — the estimates are unbiased but unstable, not wrong on average.
Detection
- VIF: $\text{VIF}_j=1/(1-R_j^2)$; VIF $>10$ flagged.
- Condition number: $\sqrt{\lambda_{\max}/\lambda_{\min}}$; $>30$ flagged.
- Pairwise correlations $|r|>0.8.$
Remedies
- Drop or combine variables.
- Ridge regression: $\hat\beta_R=(X'X+kI)^{-1}X'y.$
- Principal components regression.
🌍 Where it's used in real life
- Untangling correlated predictors (height and weight).
- Economic indicators that move together.
- Overlapping marketing channels.
- Sensors that duplicate information.
- Deciding which correlated feature to drop.
11. Heteroscedasticity
$\text{Var}(\varepsilon_i)=\sigma_i^2$ varies across $i.$
Consequences
- OLS still unbiased and consistent.
- OLS no longer BLUE; SE biased.
- $t,F$ tests invalid.
Tests
- Breusch–Pagan: regress $\hat e_i^2$ on $X$; $nR^2\sim\chi^2_{p-1}.$
- White's test: regress $\hat e_i^2$ on $X,X^2,$ cross products.
- Goldfeld–Quandt: split data by suspected variance variable, compare $F=S_2^2/S_1^2.$
Remedies
WLS, robust (White) standard errors: $\hat V(\hat\beta)=(X'X)^{-1}X'\hat\Omega X(X'X)^{-1},\,\hat\Omega=\text{diag}(\hat e_i^2).$🌍 Where it's used in real life
- Spending that varies more as income rises.
- Profit variability that grows with firm size.
- Correcting standard errors in economics.
- Financial data with changing volatility.
- Robust inference in cross-section studies.
12. Autocorrelation & Durbin–Watson
$\text{Cov}(\varepsilon_t,\varepsilon_s)\ne 0$ for $t\ne s$ — common in time series. AR(1): $\varepsilon_t=\rho\varepsilon_{t-1}+u_t.$
Consequences
OLS unbiased but inefficient; SEs biased.Durbin–Watson Statistic
$$d=\frac{\sum_{t=2}^n(\hat e_t-\hat e_{t-1})^2}{\sum_{t=1}^n\hat e_t^2}\approx 2(1-\hat\rho).$$ $d\approx 2$: no autocorrelation. $d<2$: positive autocorrelation. $d>2$: negative.Cochrane–Orcutt Procedure
Estimate $\hat\rho$ from residuals; transform variables: $y_t-\hat\rho y_{t-1}=\beta_0(1-\hat\rho)+\beta_1(x_t-\hat\rho x_{t-1})+u_t.$ Iterate.🌍 Where it's used in real life
- Correlated errors in monthly sales.
- Stock-return regressions.
- GDP and inflation time series.
- Weather-driven demand models.
- Spotting a missing trend or lag variable.
13. Logistic Regression
For binary $y\in\{0,1\}$: $P(y=1\mid x)=\pi(x)=\frac{e^{x'\beta}}{1+e^{x'\beta}}.$ Logit: $\log\frac{\pi}{1-\pi}=x'\beta.$
Estimation
By maximum likelihood (no closed form). Newton–Raphson / IRLS.Interpretation
$e^{\beta_j}$ = odds ratio for unit increase in $x_j.$Wald Test
$z=\hat\beta_j/\text{SE}(\hat\beta_j)\sim N(0,1)$ asymptotically.🌍 Where it's used in real life
- Predicting loan default (yes/no).
- Diagnosing disease present/absent.
- Predicting customer churn.
- Classifying email as spam.
- Click vs no-click on online ads.
14. Restricted Regression Estimation
Exact Restriction
$R\beta=r$ (exact, deterministic). Restricted estimator: $$\hat\beta_R=\hat\beta+(X'X)^{-1}R'[R(X'X)^{-1}R']^{-1}(r-R\hat\beta).$$ $\text{Var}(\hat\beta_R)\le\text{Var}(\hat\beta)$ in PSD ordering — variance reduction.Stochastic Restriction
$r=R\beta+v,\,v\sim N(0,\Phi)$ — Theil–Goldberger Mixed Estimator: $$\hat\beta_M=(X'X+R'\Phi^{-1}R)^{-1}(X'y+R'\Phi^{-1}r).$$Mixed Restrictions
Combine deterministic and stochastic constraints — generalized form using GLS framework.🌍 Where it's used in real life
- Imposing economic constraints like returns to scale.
- Budget shares that must sum to one.
- Using prior knowledge to stabilise estimates.
- Enforcing a known sign or size on a coefficient.
- Blending survey data with expert priors.
15. Stochastic Regressors & Errors-in-Variables (EIV)
Stochastic Regressors
If $X$ random but independent of $\varepsilon$: OLS still unbiased and consistent. If $\text{Cov}(X,\varepsilon)\ne 0$: OLS biased and inconsistent.Errors-in-Variables Model
Observe $X^*=X+u,\,Y=\beta_0+\beta_1 X+\varepsilon$ with $X$ unobserved. OLS based on $X^*$ gives $\text{plim}\,\hat\beta_1=\beta_1\frac{\sigma_X^2}{\sigma_X^2+\sigma_u^2}<\beta_1$ — attenuation bias toward 0.🌍 Where it's used in real life
- Measurement error in self-reported income.
- Test-score models with noisy inputs.
- Sensor error in engineering models.
- Bias from self-reported survey data.
- Correcting attenuation bias in econometrics.
16. Instrumental Variable (IV) Estimator
Instrument $Z$ satisfies: (i) $\text{Cov}(Z,X)\ne 0$ (relevance), (ii) $\text{Cov}(Z,\varepsilon)=0$ (exogeneity). $$\hat\beta_{IV}=(Z'X)^{-1}Z'y.$$ Consistent for $\beta$ even when $X$ is endogenous; less efficient than OLS when OLS is valid.
Intuition. When $X$ is correlated with the error, OLS cannot tell the effect of $X$ apart from the error moving with it, so it is biased. A valid instrument $Z$ shifts $X$ (relevance) only through channels unrelated to the error (exogeneity); using just that “clean” variation in $X$ recovers a consistent slope. Weak instruments (low $\text{Cov}(Z,X)$) make this cure unreliable.
🌍 Where it's used in real life
- Returns to schooling using distance to college.
- Demand estimation using weather as a supply shifter.
- Policy effects using eligibility rules.
- Health effect of a habit using taxes as an instrument.
- Causal effects when a regressor is endogenous.
17. Simultaneous Equations Model & Identification
Multiple endogenous variables determined jointly: $$Y\Gamma+XB=U,$$ $Y$: $T\times M$ endogenous, $X$: $T\times K$ exogenous.
Reduced Form
$Y=X\Pi+V,\,\Pi=-B\Gamma^{-1}.$ Always estimable by OLS.Identification (Order Condition)
For an equation with $m$ included endogenous and $k$ included exogenous variables:- Underidentified: $K-k<m-1.$
- Just identified: $K-k=m-1.$
- Overidentified: $K-k>m-1.$
Rank Condition
The matrix of excluded variables' coefficients in other equations has rank $M-1$ — necessary and sufficient.🌍 Where it's used in real life
- Supply and demand set together.
- Macro models of income and consumption.
- Wage–price systems.
- Market-equilibrium modelling.
- Interdependent policy variables.
18. 2SLS & k-class Estimator
Two-Stage Least Squares (2SLS)
- Regress each endogenous regressor on all exogenous variables → get fitted values $\hat X.$
- Replace endogenous regressor with $\hat X$ in original equation, run OLS.
k-class Estimator
Generalized form: $$\hat\beta_k=(X'(I-kM)X)^{-1}X'(I-kM)y,$$ where $M=I-Z(Z'Z)^{-1}Z'.$- $k=0$: OLS.
- $k=1$: 2SLS.
- $k$ = LIML root: Limited Information Maximum Likelihood (asymptotically equivalent to 2SLS).
🌍 Where it's used in real life
- Estimating demand with an endogenous price.
- Labour-supply and wage systems.
- Estimating macro-econometric models.
- Causal inference in economics.
- Robust estimation with weak instruments (LIML).