Skip to the content

Topics Covered

Descriptive Measures Autocovariance Function (ACVF) Autocorrelation Function (ACF) Partial ACF (PACF) Correlogram Bartlett's Formula Ergodicity General Linear Process Wold Decomposition Invertibility Stationarity MA(∞) Representation

Topic Overview — What & Why

Unit VII deals with data observed over time — ordered observations exhibit temporal dependence that classical iid methods cannot handle. Time-series analysis builds models for this dependence and uses them to forecast.

  • Time series & descriptive measures: components (trend, seasonality, cycle, irregular); sample autocovariance and autocorrelation summarise temporal dependence.
  • ACVF, ACF, PACF, correlogram: the diagnostic plots that reveal structure — cut-off vs decay patterns identify MA vs AR.
  • Strong & weak stationarity, ergodicity: stationarity ensures that statistical properties don't drift with time; ergodicity allows time-averages to estimate ensemble averages.
  • General linear process & Wold decomposition: any stationary process equals a deterministic part plus an MA($\infty$) of innovations — foundation of linear time-series modelling.
  • MA, AR, ARMA processes: three model families; stationarity (AR roots) and invertibility (MA roots) conditions must hold for usable models.
  • Yule-Walker equations: link AR coefficients to autocorrelations; provide simple estimation method.
  • Identification, estimation, order selection: Box-Jenkins methodology — identify $(p,q)$ from ACF/PACF, estimate, diagnose with Ljung-Box, select via AIC/BIC.
  • Forecasting: minimum MSE prediction using past observations; forecast variance grows with horizon.
  • Random walk & ARIMA: non-stationary cases (unit roots, integrated processes); differencing $d$ times to achieve stationarity.
  • Spectral analysis & periodogram: view the series in the frequency domain; detect hidden periodicities and energy concentration. Spectral density is the Fourier transform of the autocovariance.

1. Time Series & Descriptive Measures

Why this section? Before any model, students must learn how time-series data differ from iid samples and how to describe them with sample autocovariance and autocorrelation.

A time series is a sequence of observations $\{Y_t\}$ ordered in time. Components: trend $T_t$, seasonality $S_t$, cycle $C_t$, irregular $I_t.$ Models: additive $Y_t=T_t+S_t+C_t+I_t,$ multiplicative $Y_t=T_t\cdot S_t\cdot C_t\cdot I_t.$

Descriptive Measures

EXAMPLE 1 Quarterly retail sales: trend (long-term growth), seasonal peaks in Q4, business cycle, random noise.
EXAMPLE 2 Daily stock returns typically have $\bar Y\approx 0$ but exhibit volatility clustering — autocorrelation in $Y_t^2.$

🌍 Where it's used in real life

  1. Tracking monthly sales.
  2. Daily temperature records.
  3. A stock's price history.
  4. Website traffic over time.
  5. Electricity-demand logs.

2. Autocovariance, ACF, PACF, Correlogram

Autocovariance Function (ACVF)

$$\gamma_k=\text{Cov}(Y_t,Y_{t+k})=E[(Y_t-\mu)(Y_{t+k}-\mu)],\quad \gamma_0=\sigma_Y^2.$$

Autocorrelation Function (ACF)

$$\rho_k=\gamma_k/\gamma_0,\quad \rho_0=1,\,|\rho_k|\le 1.$$

Partial ACF (PACF)

$\phi_{kk}$ = correlation between $Y_t$ and $Y_{t-k}$ removing effect of intermediate lags. Found via Yule–Walker equations.

Correlogram

Plot of $\rho_k$ vs $k$. For white noise: $\rho_k=0$ for $k\ne 0$; $r_k\sim N(0,1/T)$ approximately. 95% bands: $\pm 1.96/\sqrt T.$

Bartlett's Formula

For MA($q$): $\text{Var}(r_k)\approx \frac{1}{T}(1+2\sum_{i=1}^q\rho_i^2)$ for $k>q.$
EXAMPLE 1 For MA(1) with $\theta=0.5$: $\rho_1=0.5/(1+0.25)=0.4,\,\rho_k=0$ for $k\ge 2.$ Correlogram cuts off after lag 1.
EXAMPLE 2 For AR(1) with $\phi=0.8$: $\rho_k=0.8^k$ — geometric decay. PACF: $\phi_{11}=0.8,\,\phi_{kk}=0$ for $k\ge 2.$
01234567AR(1): ACF tails off (ρₖ=φᵏ)lag k 01234567MA(1): ACF cuts off after lag 1lag k
Identification signature. An AR(p) has an ACF that tails off and a PACF that cuts off at lag p; an MA(q) is the mirror image — its ACF cuts off after lag q. This contrast is how Box–Jenkins reads (p,q) off the correlogram.

🌍 Where it's used in real life

  1. Detecting weekly or seasonal patterns in sales.
  2. Checking whether returns are predictable.
  3. Diagnosing the order of a forecasting model.
  4. Finding the lag structure in demand.
  5. Testing residuals for leftover pattern.

3. Stationarity & Ergodicity

Strong (strict) stationary: joint distribution of $(Y_{t_1},\ldots,Y_{t_k})$ equals that of $(Y_{t_1+h},\ldots,Y_{t_k+h})$ for all $h.$
Weak (covariance / second-order) stationary: $E(Y_t)=\mu$, $\text{Var}(Y_t)=\sigma^2<\infty,\,\gamma_k$ depends only on $k.$

Strong stationarity with finite second moments $\Rightarrow$ weak stationarity. Converse not true unless Gaussian.

Intuition. Stationarity is what makes a single observed series useful: because the statistical “rules” (mean, variance, autocorrelations) do not drift over time, we can pool information across time to estimate them. Without it, every time point would effectively be a sample of size one from a different distribution.

Ergodicity

A process is ergodic in mean if $\bar Y_T\xrightarrow{P}\mu$ as $T\to\infty.$ Sufficient: $\sum_{k=0}^\infty|\gamma_k|<\infty.$
EXAMPLE 1 White noise $\varepsilon_t\sim$ iid$(0,\sigma^2)$ — strictly stationary and ergodic.
EXAMPLE 2 $Y_t=A+\varepsilon_t$ with random $A$ but fixed for all $t$: weakly stationary but not ergodic (sample mean $\to A$, not $E(A)$).

🌍 Where it's used in real life

  1. Checking a series is stable before modelling.
  2. Detecting a trend or drift in a process.
  3. Differencing data to prepare for forecasts.
  4. Monitoring process stability.
  5. Validating that time-averages are meaningful.

4. General Linear Process & Wold Decomposition

General Linear Process

$$Y_t=\mu+\sum_{j=0}^\infty\psi_j\varepsilon_{t-j},\quad \varepsilon_t\sim\text{WN}(0,\sigma^2),\,\sum\psi_j^2<\infty.$$

Wold Decomposition

Any zero-mean weakly stationary process can be written as $$Y_t=\sum_{j=0}^\infty\psi_j\varepsilon_{t-j}+V_t,$$ where $\varepsilon_t$ is white noise (innovations), $\psi_0=1,\sum\psi_j^2<\infty$, and $V_t$ is a deterministic component (perfectly predictable from its past).
EXAMPLE 1 Pure deterministic: $Y_t=\cos(\omega t)$ — entirely predictable, $V_t=Y_t.$
EXAMPLE 2 ARMA process can be written in MA($\infty$) form: $\psi_j$ are coefficients in $\Psi(B)=\Theta(B)/\Phi(B).$

🌍 Where it's used in real life

  1. The foundation for ARMA modelling.
  2. Separating predictable from random parts.
  3. Signal-plus-noise decomposition.
  4. Understanding the limits of forecastability.
  5. Building linear forecasting models.

5. Moving Average (MA) Process

MA($q$): $$Y_t=\mu+\varepsilon_t+\theta_1\varepsilon_{t-1}+\cdots+\theta_q\varepsilon_{t-q}.$$

Properties

Invertibility

MA($q$) invertible if all roots of $\theta(B)=1+\theta_1 B+\cdots+\theta_q B^q$ lie outside unit circle. Then $\varepsilon_t=\sum\pi_j Y_{t-j}.$

Intuition. Invertibility lets us rewrite the unobservable shock $\varepsilon_t$ as a convergent weighted sum of present and past observed values $Y_{t-j}$ — without it, forecasting is impossible. It also fixes identifiability: an MA(1) with parameter $\theta$ and one with $1/\theta$ have identical ACFs, and the invertibility requirement $|\theta|<1$ singles out the unique usable version.

EXAMPLE 1 MA(1): $Y_t=\varepsilon_t+\theta\varepsilon_{t-1}.$ $\rho_1=\theta/(1+\theta^2),\,\rho_k=0$ for $k\ge 2.$ Invertible iff $|\theta|<1.$
EXAMPLE 2 MA(2): $Y_t=\varepsilon_t+0.5\varepsilon_{t-1}-0.3\varepsilon_{t-2}.$ ACF nonzero only at lags 1,2.

🌍 Where it's used in real life

  1. Smoothing short-lived shocks in demand.
  2. Modelling noisy measurement series.
  3. Economic surprise/innovation effects.
  4. Filtering short-term quality signals.
  5. Modelling short-memory data.

6. Autoregressive (AR) Process

AR($p$): $$Y_t=\phi_1 Y_{t-1}+\cdots+\phi_p Y_{t-p}+\varepsilon_t.$$

Stationarity

AR($p$) stationary iff all roots of $\phi(B)=1-\phi_1 B-\cdots-\phi_p B^p$ lie outside unit circle.

Properties

EXAMPLE 1 AR(1) $Y_t=0.7Y_{t-1}+\varepsilon_t,\,\sigma^2=1.$ $\gamma_0=1/(1-0.49)=1.96,\rho_k=0.7^k.$ Stationary since $|\phi|<1.$
EXAMPLE 2 AR(2) $Y_t=0.5Y_{t-1}+0.3Y_{t-2}+\varepsilon_t.$ Roots of $1-0.5B-0.3B^2=0$: $B=(-0.5\pm\sqrt{0.25+1.2})/0.6\approx 1.17,\,-2.84$ → both $|B|>1$, stationary.

🌍 Where it's used in real life

  1. Forecasting from recent past values.
  2. Persistence in interest rates and inflation.
  3. Predicting temperature or river flow.
  4. Modelling speech and audio.
  5. Inventory and demand forecasting.

7. ARMA Process & Conditions

ARMA($p,q$): $$\phi(B)Y_t=\theta(B)\varepsilon_t.$$

MA($\infty$) Representation

$Y_t=\Psi(B)\varepsilon_t,\,\Psi(B)=\theta(B)/\phi(B).$

AR($\infty$) Representation

$\Pi(B)Y_t=\varepsilon_t,\,\Pi(B)=\phi(B)/\theta(B).$
EXAMPLE 1 ARMA(1,1) $Y_t=0.5Y_{t-1}+\varepsilon_t+0.3\varepsilon_{t-1}.$ Stationary ($|0.5|<1$), invertible ($|0.3|<1$).
EXAMPLE 2 For ARMA(1,1) above: $\rho_1=\frac{(1+\phi\theta)(\phi+\theta)}{1+2\phi\theta+\theta^2}=\frac{(1.15)(0.8)}{1.39}\approx 0.662,\,\rho_k=0.5\rho_{k-1}$ for $k\ge 2.$

🌍 Where it's used in real life

  1. General short-term sales and demand forecasts.
  2. Modelling economic indicators.
  3. Energy-load forecasting.
  4. Modelling correlated process data.
  5. Signal modelling in engineering.

8. Yule–Walker Equations

For AR($p$), multiplying by $Y_{t-k}$ and taking expectations: $$\rho_k=\phi_1\rho_{k-1}+\cdots+\phi_p\rho_{k-p},\quad k\ge 1.$$

Matrix form for first $p$ equations: $$\begin{pmatrix}\rho_1\\\vdots\\\rho_p\end{pmatrix}=\begin{pmatrix}1&\rho_1&\cdots&\rho_{p-1}\\\rho_1&1&\cdots&\rho_{p-2}\\\vdots&&\ddots&\\\rho_{p-1}&\cdots&\rho_1&1\end{pmatrix}\begin{pmatrix}\phi_1\\\vdots\\\phi_p\end{pmatrix}.$$ Solve to obtain Yule–Walker estimates of $\phi_j.$

EXAMPLE 1 AR(1): $\rho_1=\phi_1\Rightarrow\hat\phi_1=r_1.$ Then $\hat\phi_k=r_1^k$ predicted ACF.
EXAMPLE 2 AR(2) with $r_1=0.6,r_2=0.4.$ Solve $\rho_1=\phi_1+\phi_2 r_1,\,\rho_2=\phi_1 r_1+\phi_2.$ $\Rightarrow$ $\phi_1=\frac{\rho_1(1-\rho_2)}{1-\rho_1^2}=0.5625,\ \phi_2=\frac{\rho_2-\rho_1^2}{1-\rho_1^2}=0.0625.$

🌍 Where it's used in real life

  1. Estimating AR coefficients for forecasts.
  2. Linear prediction in speech coding.
  3. Spectral estimation.
  4. Fitting autoregressive demand models.
  5. Radar and signal parameter estimation.

9. Identification, Estimation, Order Selection

ProcessACFPACF
White noise0 for all $k\ne 0$0 for all $k\ne 0$
AR($p$)tails offcuts off after lag $p$
MA($q$)cuts off after lag $q$tails off
ARMA($p,q$)tails offtails off

Estimation

Order Selection Criteria

$$\text{AIC}=-2\log L+2k,\qquad \text{BIC}=-2\log L+k\log T.$$ Minimize over $(p,q)$. BIC penalizes more heavily for large $T$ — chooses parsimonious models.

Box–Jenkins Methodology

  1. Identification: examine ACF/PACF.
  2. Estimation.
  3. Diagnostic checking: residual ACF, Ljung–Box test $Q^*=T(T+2)\sum_{k=1}^h r_k^2/(T-k)\sim\chi^2_{h-p-q}.$
  4. Forecasting.
EXAMPLE 1 Sample ACF cuts off at lag 2, PACF tails off → MA(2).
EXAMPLE 2 AIC values: AR(1)=120, AR(2)=115, AR(3)=116. Choose AR(2).

🌍 Where it's used in real life

  1. Choosing the right model with AIC/BIC.
  2. Box–Jenkins model building for sales.
  3. Avoiding over- and under-fitting.
  4. Residual diagnostics (Ljung–Box).
  5. Automated forecasting pipelines.

10. Forecasting

Minimum MSE forecast for $h$-step ahead: $\hat Y_t(h)=E(Y_{t+h}\mid \mathcal F_t).$ For ARMA in MA($\infty$) form: $$\hat Y_t(h)=\sum_{j=h}^\infty\psi_j\varepsilon_{t-j+h}.$$ Forecast error: $e_t(h)=\sum_{j=0}^{h-1}\psi_j\varepsilon_{t+h-j},\,V[e_t(h)]=\sigma^2\sum_{j=0}^{h-1}\psi_j^2.$

AR(1) Forecast

$\hat Y_t(h)=\phi^h Y_t.$ Variance $\sigma^2(1-\phi^{2h})/(1-\phi^2).$

Exponential Smoothing (Holt–Winters)

Optimal under specific ARIMA: $\text{ETS}(A,N,N)\equiv\text{ARIMA}(0,1,1).$
EXAMPLE 1 AR(1) $\phi=0.6,\,Y_T=10.$ Forecasts: $\hat Y(1)=6,\hat Y(2)=3.6,\,\to 0.$
EXAMPLE 2 MA(1): $\hat Y_t(1)=\theta\varepsilon_t,\,\hat Y_t(h)=0$ for $h\ge 2.$ Variance: $\sigma^2$ for $h=1$, $\sigma^2(1+\theta^2)$ for $h\ge 2.$

🌍 Where it's used in real life

  1. Next-quarter sales forecasts.
  2. Weather and rainfall prediction.
  3. Planning electricity demand.
  4. Inventory planning from demand forecasts.
  5. Staffing forecasts for call centres.

11. Non-stationary Time Series — Random Walk & ARIMA

Random Walk

$Y_t=Y_{t-1}+\varepsilon_t.$ Non-stationary: $\text{Var}(Y_t)=t\sigma^2\to\infty.$ With drift: $Y_t=\delta+Y_{t-1}+\varepsilon_t.$ First differences $\Delta Y_t=\varepsilon_t$ are stationary.

ARIMA($p,d,q$)

$\phi(B)(1-B)^d Y_t=\theta(B)\varepsilon_t.$ Differencing $d$ times yields a stationary ARMA process.

Tests for Unit Root

Parameter Estimation

After differencing, fit ARMA via MLE / least squares. Include constant if drift suspected.
EXAMPLE 1 Stock prices typically modelled as ARIMA(0,1,0) — random walk; returns are white noise.
EXAMPLE 2 ARIMA(1,1,1): $(1-\phi B)(1-B)Y_t=(1+\theta B)\varepsilon_t.$ First-difference, fit ARMA(1,1) on $\Delta Y.$

🌍 Where it's used in real life

  1. Modelling stock prices as a random walk.
  2. Forecasting GDP and inflation by differencing.
  3. Sales with a trend via ARIMA.
  4. Modelling exchange rates.
  5. Testing for unit roots in economics.

12. Spectral Analysis

Spectral Density

For stationary process with absolutely summable ACVF: $$f(\omega)=\frac{1}{2\pi}\sum_{k=-\infty}^\infty\gamma_k e^{-ik\omega},\quad\omega\in[-\pi,\pi].$$ Inverse: $\gamma_k=\int_{-\pi}^\pi f(\omega)e^{ik\omega}d\omega.$ $f(\omega)\ge 0,\,f(-\omega)=f(\omega).$

White noise

$f(\omega)=\sigma^2/(2\pi)$ — flat.

AR(1)

$f(\omega)=\frac{\sigma^2}{2\pi|1-\phi e^{-i\omega}|^2}=\frac{\sigma^2}{2\pi(1-2\phi\cos\omega+\phi^2)}.$

MA(1)

$f(\omega)=\frac{\sigma^2}{2\pi}|1+\theta e^{-i\omega}|^2=\frac{\sigma^2}{2\pi}(1+2\theta\cos\omega+\theta^2).$

ARMA

$f(\omega)=\frac{\sigma^2}{2\pi}\frac{|\theta(e^{-i\omega})|^2}{|\phi(e^{-i\omega})|^2}.$
EXAMPLE 1 AR(1) with $\phi=0.9$: spectral density large at low frequencies — long-run cycles.
EXAMPLE 2 AR(1) with $\phi=-0.9$: density large at high frequencies — alternating pattern.

🌍 Where it's used in real life

  1. Finding cycles and periodicities in data.
  2. Vibration analysis of machinery.
  3. EEG and ECG frequency analysis.
  4. Detecting business cycles.
  5. Frequency content of audio signals.

13. Periodogram & Spectral Estimation

Periodogram

$$I(\omega_j)=\frac{1}{2\pi T}\left|\sum_{t=1}^T Y_t e^{-i\omega_j t}\right|^2,\quad \omega_j=2\pi j/T.$$ $E[I(\omega)]\to f(\omega)$ as $T\to\infty$, but $\text{Var}[I(\omega)]\not\to 0$ — inconsistent.

Smoothed Periodogram

$\hat f(\omega)=\sum_k W_k I(\omega_{j+k})$ — average periodogram values over neighbouring frequencies. Window functions: Daniell, Bartlett, Parzen.

Frequency Resolution vs Variance Trade-off

Wider window → lower variance but lower resolution.
EXAMPLE 1 For sinusoidal series $Y_t=A\cos(\omega_0 t+\phi)+\varepsilon_t$: periodogram has peak at $\omega_0.$ Used for hidden periodicity detection.
EXAMPLE 2 Fisher's $g$-test: under white noise null, max ordinate ratio $g=I_{\max}/\sum I_j$ has known distribution.

🌍 Where it's used in real life

  1. Detecting hidden seasonal cycles.
  2. Sunspot and climate cycle detection.
  3. Finding machinery fault frequencies.
  4. Tidal and astronomical periodicities.
  5. Signal periodicity in engineering.