Econometrics (literally, "measurement of economic things") is the branch of economics
that uses statistical and mathematical methods to test economic theories and to estimate economic
relationships from real-world data. Coined by Ragnar Frisch in 1933, it sits at the intersection of
three disciplines:
Economic theory — supplies the hypotheses (e.g. demand decreases with price).
Mathematics — gives precise functional forms (e.g. \(Q = a + bP\)).
Statistics — provides estimation and inference (e.g. estimate \(b\), test \(b < 0\)).
1.1 How Econometrics Differs From "Plain" Statistics
Statistics
Econometrics
General methodology, no specific subject matter.
Built around economic theories.
Most data are experimental (controlled).
Data are observational; we cannot run controlled experiments.
Assumes well-behaved error structure.
Has to deal with heteroscedasticity, autocorrelation, multicollinearity etc.
Causal interpretation is rarely required.
Causal interpretation is central.
2. Steps in an Empirical Economic Analysis
Every econometric study follows the same broad pipeline:
Statement of theory or hypothesis — e.g. "Income causes consumption."
Mathematical model — \(C = \beta_0 + \beta_1 Y\) (deterministic).
Econometric model — \(C_i = \beta_0 + \beta_1 Y_i + u_i\) (with a stochastic
disturbance \(u_i\) capturing everything not modelled).
Data collection — household income/consumption surveys.
Estimation of parameters — OLS gives \(\hat\beta_0, \hat\beta_1\).
Hypothesis testing — is \(\beta_1\) significantly different from zero?
Forecasting / prediction — given a forecast of income, predict consumption.
Use the model for policy — government tax cut increases income by ₹10,000;
predicted rise in aggregate consumption.
EXAMPLE 1 — Keynesian consumption function
Theory: Consumption depends positively on disposable income, but \(0 < \text{MPC} < 1\).
Econometric model: \(C_i = \beta_0 + \beta_1 Y_i + u_i\). With Indian household survey data
(NSSO), suppose OLS gives \(\hat\beta_0 = 1{,}500\) and \(\hat\beta_1 = 0.72\). Test
\(H_0: \beta_1 = 0\). If the \(t\)-statistic is highly significant we conclude income drives
consumption, with the marginal propensity to consume around 0.72.
EXAMPLE 2 — Demand for a good
Theory: Quantity demanded falls as price rises.
Mathematical model: \(\log Q = \beta_0 + \beta_1 \log P\), so \(\beta_1\) is the price elasticity.
Econometric model: \(\log Q_i = \beta_0 + \beta_1 \log P_i + u_i\). With market data, suppose
\(\hat\beta_1 = -0.85\) — price elasticity of demand is \(-0.85\), implying demand is relatively
inelastic but downward-sloping, consistent with theory.
3. The Econometric Model
GENERAL FORM
Any econometric model can be written as
\[
Y = f(X_1, X_2, \ldots, X_k; \boldsymbol\beta) + u,
\]
where \(Y\) is the dependent variable, the \(X\)'s are explanatory (independent) variables,
\(\boldsymbol\beta\) is a vector of parameters, and \(u\) is the stochastic disturbance.
3.1 Why a Stochastic Term \(u\)?
Omitted variables — many causes of \(Y\) cannot be measured (taste, culture, expectations).
Measurement error in both \(Y\) and the \(X\)s.
Inherent randomness — human decisions are not deterministic.
Wrong functional form — the true relationship may not be linear.
Aggregation errors — relationships at micro level smoothed at macro level.
These are the building blocks of the Classical Linear Regression Model (CLRM) studied in
Unit 2 onward.
4. Importance of Measurement in Economics
Without measurement, economic theory is just a story. Measurement matters because it:
Quantifies economic relationships — not just "demand falls with price" but
"demand falls by 0.85% for every 1% rise in price."
Tests economic theories — the data may reject what theory predicts.
Forecasts future values used in budgeting, planning, monetary policy.
Compares alternative policies — Will a GST rate cut increase tax revenue?
Establishes causality — distinguishing correlation from causation is the
core econometric task.
5. The Structure of Econometric Data
Four main data structures appear in empirical economics, each with its own methods:
5.1 Cross-Section Data
DEFINITION
Many units (households, firms, countries) observed at one point in time. Typically
obtained from surveys.
Example: Income, education and consumption of 5,000 Indian households in
January 2024.
Key issue: heterogeneity across units (heteroscedasticity is common).
5.2 Pooled Cross-Section Data
DEFINITION
Cross-section samples from two or more different time periods stacked together — but
not the same units in each period.
Example: NSSO consumption surveys from 2019 and 2024 (different households in
each round). Useful for studying changes in cross-sectional distributions over time, often after
a policy intervention.
5.3 Time-Series Data
DEFINITION
One unit observed at regular intervals over time — annual GDP, monthly inflation,
quarterly profits.
Example: Quarterly Indian GDP from 1991 Q1 to 2024 Q4.
Key issue: serial correlation, non-stationarity, seasonality — handled by
autocorrelation theory (Unit 5) and time-series methods.
5.4 Paired (Panel / Longitudinal) Data
DEFINITION
The same units followed across multiple time periods. Combines features of cross-section
and time-series — also called "panel data."
Example: Income of 1,000 Indian households tracked annually from 2015 to 2024
(10 years × 1,000 households = 10,000 observations).
Strength: controls for unobserved unit-specific effects (e.g. innate ability).
5.5 Quick-Reference Table
Data Type
Cross-Section
Pooled CS
Time-Series
Panel
Many units?
Yes
Yes
1 (usually)
Yes
Many periods?
1
2+
Many
Many
Same units across periods?
—
No
—
Yes
Typical problem
Heteroscedasticity
Both
Autocorrelation
Both + unobserved effects
EXAMPLE 1 — Identifying data structures
The CMIE consumer pyramids household survey, December 2023: cross-section.
The same households surveyed in 2018, 2020 and 2023: panel.
RBI monthly WPI inflation series 2000–2024: time-series.
NSSO unemployment surveys (different households) in 2015 and 2024: pooled cross-section.
EXAMPLE 2 — Why structure matters
Suppose we want to study the effect of education on wages.
With cross-section data we cannot rule out that "unobserved ability" causes
both more education and higher wages — biased estimate.
With panel data we can take "within-person" differences: if the same person
earns more after their education increases, that effect is identified — much cleaner.
Same theory; different conclusions depending on data structure.
6. The Three Roles of Econometrics
Empirical verification of theory — does Keynes's MPC really lie in (0,1)?
Policy analysis — what is the impact of a 1% repo rate cut on housing demand?
Forecasting — what will GDP growth be next quarter?
Cautionary note: Econometrics can refute theories but cannot prove
them. A statistically significant fit is not the same as causation; correlation between two series
might reflect a third variable or pure coincidence. Always pair quantitative results with economic
judgement.
Key Take-aways from Unit 1
Econometrics = Economic theory + Mathematics + Statistics + Data.
Empirical analysis follows a fixed pipeline: theory → model → data → estimate → test → use.
The econometric model contains a stochastic disturbance \(u\) to capture omitted variables,
measurement error and randomness.
Four data structures: cross-section, pooled cross-section, time-series, panel.
Cross-section data tends to bring heteroscedasticity; time-series brings autocorrelation;
panel data brings both plus unobserved effects.