Part of the data-platform path: Big Data Technologies, Cloud Computing for Data Science, Time Series Analysis and Forecasting and Data Engineering and MLOps.
Everything you learned in Statistical Foundations for Data Science, Data Mining and Machine Learning assumed your observations were independent. In a time series they are not, and that single fact breaks almost every technique you know.
| What you learned | Why it fails here |
|---|---|
| Shuffle, then split 80/20 | A random split lets the model see the future |
| k-fold cross-validation | Every fold trains on data that comes after the validation fold |
| The standard error of the mean | Assumes independent draws; correlated data has far less information than n suggests |
| Bootstrap resampling | Destroys the ordering, which is the signal |
| "More data is more information" | 400 correlated points carry the information of far fewer |
The ordering is not metadata. It is the data, and every method in this course exists to model dependence between an observation and the ones before it.
THE BIG IDEA
Stationarity. A stationary series has a constant mean, a constant variance, and an autocovariance that depends only on the gap between two points and not on when they are.
Almost every model here requires it. Units 1 and 3 are about testing for it and getting there, and once you have it, the modelling in Unit 2 is comparatively easy.
Everything. This is one of only two courses in the whole catalogue with no
NOT EXECUTED file anywhere — the other is Machine Learning.
statsmodels 0.14 implements every technique the syllabus names:
| Technique | Real call |
|---|---|
| Decomposition, STL | seasonal_decompose, STL |
| ACF, PACF | acf, pacf |
| Stationarity | adfuller, kpss |
| ARMA, ARIMA, SARIMA | ARIMA, SARIMAX |
| Diagnostics | acorr_ljungbox |
| Multivariate | VAR, test_causality |
| State space | UnobservedComponents (Kalman filter) |
| Spectral | scipy.signal.periodogram |
| Exponential smoothing | ExponentialSmoothing |
pip install -r tools/requirements.txt
python3 tools/data-science/run_timeseries_labs.py
KEY INSIGHT
The series are generated from KNOWN coefficients, so every fit can be checked against the truth that produced it:
A model that merely converges has proved nothing. Checking against a known answer is the only way to show that identification and estimation work, and it is why the labs found — and report — a coefficient estimate 2.5 standard errors from the truth on one draw.
The course aims to:
Provide fundamental understanding of time series data, components, and characteristics.
Train students in identifying, modeling, and forecasting using ARMA/ARIMA/SARIMA models.
Introduce state-space and multivariate approaches for complex data.
Familiarize students with modern forecasting methods, including spectral and evaluation techniques.
Enable hands-on practice with real-world datasets using R/Python statistical libraries.
WATCH OUT
The evaluation method names both languages, and the practicals assume whichever you have. These notes use Python throughout and say so; a student following the syllabus literally has no basis for the choice. See review finding D30.
| Unit | Question it answers |
|---|---|
| 1 | What is a time series, and is this one stationary? |
| 2 | How do I model a stationary series, and forecast it? |
| 3 | What if it is not stationary, or it has a season? |
| 4 | What if there are several series, or the state is hidden? |
| 5 | Which cycles is it made of, and which forecast is best? |
Units 1–3 are the spine. Unit 4 is two genuinely different frameworks bolted on, and Unit 5 is where you learn that the answer to "which model is best?" depends on a metric you should have chosen first.
labs/course-14b-timeseries/ — the code, and the runner that asserts every figure
these notes quote
data/course-14b-timeseries/ — practice datasets, CSV: ar2-series.csv, macro-indicators.csv, seasonal-sales.csv.
Every one was generated from a known truth, so you can score your answer
rather than just produce one; data/README.md lists what each was built
from, data/PRACTICE-QUESTIONS.md sets questions on each with a computed
answer key, and tools/data-science/check_datasets.py proves every one of those
answers against the file.
| From | To | What is shared |
|---|---|---|
| Statistical Foundations for Data Science (Statistics) | Units 1, 2 | Hypothesis tests, p-values and confidence intervals — used constantly, and the ADF's reversed null is where they bite |
| Machine Learning (ML) | Unit 5 | The bias–variance trade-off, arriving as AIC against held-out RMSE — and a train/test split you must do differently |
| Python for Data Analysis and Visualization (Pandas) | throughout | Every series here is a pandas object; resample, shift and rolling are the working tools |
| Data Science with R | the lab | The syllabus says "R/Python". R's forecast package is the reference implementation; statsmodels is what runs here |
| Data Engineering and MLOps | Unit 5 | Rolling-origin backtesting is what monitoring a deployed forecast actually means |
References: Box, Jenkins & Reinsel, Time Series Analysis: Forecasting and Control — the origin of the ARIMA notation this course uses · Montgomery, Jennings & Kulahci, Introduction to Time Series Analysis and Forecasting, Wiley · Shumway & Stoffer, Time Series Analysis and Its Applications: With R Examples.
Free and current: Hyndman & Athanasopoulos, Forecasting: Principles and
Practice, is open access at otexts.com/fpp3. It is
not on the syllabus, and it is the book most working forecasters actually use.
Learn the two identification rules and be able to apply them. ACF cuts off → MA(q); PACF cuts off → AR(p). Everything in Unit 2 hangs on them.
Draw an ACF by hand for white noise, an AR(1) with φ = 0.8, and an MA(1). If you can sketch those three, you can read any correlogram.
Know the ADF's null by heart. It is the reverse of what you expect, and getting it backwards inverts every conclusion you draw.
Do the differencing arithmetic. d for trend, D for season, and stop as soon as the ADF rejects.
Run the labs. Thirteen experiments, all executable, all checked against generated truth.
"The ADF's null is that a unit root is present, so a small p-value means stationary."
"The PACF cuts off at lag p for an AR(p); the ACF cuts off at lag q for an MA(q)."
"k-fold cross-validation is invalid on a time series; use rolling origin."
"Granger causality means 'helps predict', not 'causes'."
Why a random split lets a model see the future, and which of your habits from Courses 4, 8 and 12 A stop being valid; time series types, components and the forecasting process; stationarity defined; autocovariance, ACF and PACF, and reading them together to identify a model; decomposition into trend, seasonal and residual, classical against STL, checked against known coefficients.
UNIT 2AR, MA and ARMA(p,q) definitions and their signatures in the ACF and PACF; estimation, and a 200-draw Monte Carlo showing the estimator unbiased when a single fit misses; AIC and BIC against rolling-origin cross-validation, and what to do when they disagree; Ljung–Box residual diagnostics; forecasting from an ARMA and where the intervals come from.
UNIT 3Differencing, the ADF and KPSS tests, and using the ct regression to tell a trend-stationary series from a random walk; over-differencing measured as a variance increase; ARIMA and SARIMA with seasonal orders; the airline model; prediction intervals and a measured coverage of 75% against a nominal 95%.
UNIT 4Vector autoregression on a macro system built with known causality; Granger causality tested in all four directions; state-space form and the Kalman filter through unobserved components; what a zero variance estimate means; spectral analysis and the periodogram, and why detrending first changes the answer.
UNIT 5ARIMA against exponential smoothing against machine learning on the same series; naive and seasonal-naive baselines; RMSE, MAE, MAPE and MASE, and the cases where they disagree; why MAPE reaches 75% on a uniform error of 1.0; recursive forecasting and a measured result that contradicts the textbook; tree models that cannot extrapolate; testing a forecast for bias.
PRACTICEExam-style questions with fully worked solutions.
LABEvery prescribed lab experiment, with code and expected output.