Skip to the content

Topics Covered

Nature of Econometrics Concept Steps in Empirical Analysis Econometric Model Measurement Cross-Section Pooled Cross-Section Time-Series Paired Data
On this page
  1. 1. What is Econometrics?
  2. 2. Steps in an Empirical Economic Analysis
  3. 3. The Econometric Model
  4. 4. Importance of Measurement in Economics
  5. 5. The Structure of Econometric Data
  6. 6. The Three Roles of Econometrics
  7. Key Take-aways from Unit 1

1. What is Econometrics?

DEFINITION

Econometrics (literally, "measurement of economic things") is the branch of economics that uses statistical and mathematical methods to test economic theories and to estimate economic relationships from real-world data. Coined by Ragnar Frisch in 1933, it sits at the intersection of three disciplines:

1.1 How Econometrics Differs From "Plain" Statistics

StatisticsEconometrics
General methodology, no specific subject matter.Built around economic theories.
Most data are experimental (controlled).Data are observational; we cannot run controlled experiments.
Assumes well-behaved error structure.Has to deal with heteroscedasticity, autocorrelation, multicollinearity etc.
Causal interpretation is rarely required.Causal interpretation is central.

2. Steps in an Empirical Economic Analysis

Every econometric study follows the same broad pipeline:

  1. Statement of theory or hypothesis — e.g. "Income causes consumption."
  2. Mathematical model — \(C = \beta_0 + \beta_1 Y\) (deterministic).
  3. Econometric model — \(C_i = \beta_0 + \beta_1 Y_i + u_i\) (with a stochastic disturbance \(u_i\) capturing everything not modelled).
  4. Data collection — household income/consumption surveys.
  5. Estimation of parameters — OLS gives \(\hat\beta_0, \hat\beta_1\).
  6. Hypothesis testing — is \(\beta_1\) significantly different from zero?
  7. Forecasting / prediction — given a forecast of income, predict consumption.
  8. Use the model for policy — government tax cut increases income by ₹10,000; predicted rise in aggregate consumption.
EXAMPLE 1 — Keynesian consumption function

Theory: Consumption depends positively on disposable income, but \(0 < \text{MPC} < 1\).

Mathematical model: \(C = \beta_0 + \beta_1 Y\) with \(0 < \beta_1 < 1\).

Econometric model: \(C_i = \beta_0 + \beta_1 Y_i + u_i\). With Indian household survey data (NSSO), suppose OLS gives \(\hat\beta_0 = 1{,}500\) and \(\hat\beta_1 = 0.72\). Test \(H_0: \beta_1 = 0\). If the \(t\)-statistic is highly significant we conclude income drives consumption, with the marginal propensity to consume around 0.72.

EXAMPLE 2 — Demand for a good

Theory: Quantity demanded falls as price rises.

Mathematical model: \(\log Q = \beta_0 + \beta_1 \log P\), so \(\beta_1\) is the price elasticity.

Econometric model: \(\log Q_i = \beta_0 + \beta_1 \log P_i + u_i\). With market data, suppose \(\hat\beta_1 = -0.85\) — price elasticity of demand is \(-0.85\), implying demand is relatively inelastic but downward-sloping, consistent with theory.

3. The Econometric Model

GENERAL FORM

Any econometric model can be written as

\[ Y = f(X_1, X_2, \ldots, X_k; \boldsymbol\beta) + u, \]

where \(Y\) is the dependent variable, the \(X\)'s are explanatory (independent) variables, \(\boldsymbol\beta\) is a vector of parameters, and \(u\) is the stochastic disturbance.

3.1 Why a Stochastic Term \(u\)?

3.2 Classical Assumptions on \(u\)

\[ E(u_i) = 0, \quad \mathrm{Var}(u_i) = \sigma^2, \quad \mathrm{Cov}(u_i, u_j) = 0\ (i \ne j), \] \[ \mathrm{Cov}(u_i, X_i) = 0, \quad u_i \sim N(0, \sigma^2). \]

These are the building blocks of the Classical Linear Regression Model (CLRM) studied in Unit 2 onward.

4. Importance of Measurement in Economics

Without measurement, economic theory is just a story. Measurement matters because it:

5. The Structure of Econometric Data

Four main data structures appear in empirical economics, each with its own methods:

5.1 Cross-Section Data

DEFINITION

Many units (households, firms, countries) observed at one point in time. Typically obtained from surveys.

Example: Income, education and consumption of 5,000 Indian households in January 2024.

Key issue: heterogeneity across units (heteroscedasticity is common).

5.2 Pooled Cross-Section Data

DEFINITION

Cross-section samples from two or more different time periods stacked together — but not the same units in each period.

Example: NSSO consumption surveys from 2019 and 2024 (different households in each round). Useful for studying changes in cross-sectional distributions over time, often after a policy intervention.

5.3 Time-Series Data

DEFINITION

One unit observed at regular intervals over time — annual GDP, monthly inflation, quarterly profits.

Example: Quarterly Indian GDP from 1991 Q1 to 2024 Q4.

Key issue: serial correlation, non-stationarity, seasonality — handled by autocorrelation theory (Unit 5) and time-series methods.

5.4 Paired (Panel / Longitudinal) Data

DEFINITION

The same units followed across multiple time periods. Combines features of cross-section and time-series — also called "panel data."

Example: Income of 1,000 Indian households tracked annually from 2015 to 2024 (10 years × 1,000 households = 10,000 observations).

Strength: controls for unobserved unit-specific effects (e.g. innate ability).

5.5 Quick-Reference Table

Data TypeCross-SectionPooled CSTime-SeriesPanel
Many units?YesYes1 (usually)Yes
Many periods?12+ManyMany
Same units across periods?—No—Yes
Typical problemHeteroscedasticityBothAutocorrelationBoth + unobserved effects
EXAMPLE 1 — Identifying data structures
  1. The CMIE consumer pyramids household survey, December 2023: cross-section.
  2. The same households surveyed in 2018, 2020 and 2023: panel.
  3. RBI monthly WPI inflation series 2000–2024: time-series.
  4. NSSO unemployment surveys (different households) in 2015 and 2024: pooled cross-section.
EXAMPLE 2 — Why structure matters

Suppose we want to study the effect of education on wages.

6. The Three Roles of Econometrics

  1. Empirical verification of theory — does Keynes's MPC really lie in (0,1)?
  2. Policy analysis — what is the impact of a 1% repo rate cut on housing demand?
  3. Forecasting — what will GDP growth be next quarter?
Cautionary note: Econometrics can refute theories but cannot prove them. A statistically significant fit is not the same as causation; correlation between two series might reflect a third variable or pure coincidence. Always pair quantitative results with economic judgement.

Key Take-aways from Unit 1