Skip to the content
On this page
  1. The one thing to understand before anything else
  2. What runs here
  3. Course objectives (verbatim)
  4. The five units
  5. Also here
  6. The result that surprised the lab
  7. How this course connects to the rest of the catalogue
  8. Textbooks
  9. How to study this course
  10. If you read one thing
  11. Units in this Course

Part of the data-platform path: Big Data Technologies, Cloud Computing for Data Science, Time Series Analysis and Forecasting and Data Engineering and MLOps.


The one thing to understand before anything else

Every other course in this catalogue ends when the model works. This one starts there.

The rest of the catalogue asks This course asks
does the model fit? can somebody else reproduce it?
what is the test accuracy? is it still that accurate six months from now?
which algorithm is best? what happens at 3 a.m. when it stops responding?
— can you delete one person's data from it?
— who is accountable when it is wrong?

THE BIG IDEA

The single most examinable idea

A model in production is a system, not an artefact. It has inputs that change, dependencies that drift, users who complain, regulators who ask, and an owner who must be nameable. The model file is the smallest part of it.

⚠️ The measurement that makes the point

Experiment 3 runs an ETL job on deliberately messy data. Three defects — a region spelled in lowercase, a price written with a currency prefix, a duplicated order — produce errors of −₹1,680, −₹700 and +₹2,520, which nearly cancel. The uncleaned total comes out at ₹10,500 against a true ₹10,360: wrong by about one percent.

NOTE

That is the dangerous case. A figure that is wildly wrong gets noticed; a figure that is 1% wrong gets reported to the board. The errors cancelled by luck, and next month they will not.

The only reliable check is an independently computed figure, which is why this repository has five engines agreeing on that number.


What runs here

Eleven of the sixteen experiments run against the real tools, not descriptions of them:

Tool What is real here
MLflow 3 six runs logged to a SQLite backend and queried back, ordered by AUC
git + DVC two data versions committed, dvc checkout restoring the earlier one, verified by comparing the recovered column
Flask a server on a real socket, called over HTTP, including both error paths
SQLite real constraints, which reject the three bad inserts the lab attempts
scipy KS test and PSI, scored against drift injected at a known magnitude

The five that do not run, and why

# Experiment Reason Its runnable half
4 Kafka / RabbitMQ needs a broker process 04_batch_vs_event.py — both modes over a real queue, latency measured
5 HDFS needs a JVM and a NameNode the ETL job, plus Big Data Technologies' block arithmetic
10 Docker the client is installed; the daemon is not the Flask app the container would package, running
11 GitHub Actions needs a GitHub runner the determinism check CI exists to protect
15 Prometheus / Grafana both are server processes a real /metrics endpoint, parsed back as valid exposition format

None of those halves is filler. In each case the part that this environment blocks is the infrastructure, and the part you actually write is verified.

tools/data-science/run_mlops_labs.py asserts all five *** NOT EXECUTED *** markers are still present.

IN DEPTH

Why the data is generated

You cannot verify a drift detector on data whose drift you do not know.

So the loan dataset here is built from known coefficients, and the drift is injected at a known magnitude on a known feature at a known time. That makes two things checkable that otherwise could only be reported:

Claim How it is checked
the model fitted correctly its coefficients are compared against the ones that generated the data
the drift detector works 4 of 5 drifted batches caught, 0 false alarms, with a one-batch lag

Course objectives (verbatim)

  1. To introduce the lifecycle and roles in Data Engineering.
  2. To explore data architecture principles, distributed systems, and technology choices.

  3. To analyze MLOps features, risks, and challenges in developing ML systems.

  4. To design CI/CD pipelines and deployment strategies for ML models.
  5. To understand monitoring, governance, and Responsible AI compliance in production ML.

The five units

Unit Topic Notes Hardest part
1 Foundations of data engineering unit-1.md the data lifecycle vs the data engineering lifecycle
2 Data architecture and distributed systems unit-2.md when microservices are wrong
3 MLOps fundamentals unit-3.md the four things that must be pinned
4 Deployment and CI/CD unit-4.md the metric gate, and canary vs shadow
5 Monitoring, feedback loops and governance unit-5.md why data drift is not concept drift

Plus lab.md and practice.md.


Also here

The result that surprised the lab

Experiment 14 builds an automatic retraining loop: detect drift, retrain, deploy. It fires at four batches — and improves accuracy by +0.0016, which is nothing.

NOTE

The reason is the distinction the whole of Unit 5 turns on. The drift shifted P(X) — incomes moved up. It did not shift P(y|X) — the relationship between income and approval was unchanged. A model that learned the true relationship is still correct on shifted inputs.

So: alert on data drift, investigate, and retrain only when the labels confirm the relationship moved. Retraining on every input shift is expensive and can make things worse.

That is a more useful lesson than a demo where retraining rescues the model, and it is the honest result of the code as written.


How this course connects to the rest of the catalogue

Course What it gives you here
Database Management Systems (DBMS) schemas, keys, constraints — the warehouse in experiment 3
Python for Data Analysis and Visualization (Python for Data Analysis) the pipeline and the profiling
Machine Learning (Machine Learning) the model being deployed, and the baselines
Big Data Technologies (Big Data) HDFS, Kafka, the distributed layer
Cloud Computing for Data Science (Cloud Computing) where all of this is deployed, and what it costs
Time Series Analysis and Forecasting (Time Series) drift is distribution change over time

Cross-check: the South-region revenue total of ₹10,360 is computed here by a fifth independent engine, after Business Intelligence Tools' DAX, Big Data Technologies' Hive and Spark, and Cloud Computing for Data Science's warehouse.


Textbooks

The syllabus prescribes a single combined Text / Reference list, and it has exactly one book in it:

Web resources named in the syllabus: IBM's data-engineering topic page · Martin Fowler on microservices · a Towards Data Science introduction to MLOps.

WATCH OUT

⚠️ One incomplete citation, and one resource behind a paywall

The reading list ends "Fundamentals of Data Engineering, Joe Reis & Matt Housley," — the trailing comma is where the publisher and year should be, and item 2 turns out to be the heading "Web Resources" rather than a book. This is the only course of the later ones whose students cannot locate their single prescribed text from the syllabus alone. See review finding D31.

Towards Data Science moved to a Medium members-only model, so the third web resource may be paywalled. See review finding D32. Google's Practitioners Guide to MLOps and Microsoft's MLOps documentation cover the same ground and are free.

How to study this course

  1. Run MLflow locally in week 3. mlflow ui against a SQLite backend takes five minutes and makes Unit 3 concrete. Reading about experiment tracking teaches nothing; losing a good result because you did not track it teaches it permanently.

  2. Put one real project under version control, data included. DVC's whole idea — the pointer is in git, the bytes are not — only lands once you have seen a repository stay small while the dataset changes.

  3. Learn the lifecycle as a sequence you can name. Units 1 and 2 are vocabulary questions: ingestion, storage, transformation, serving, and the architecture choices behind each. They are the easy marks.

  4. Build the smallest possible CI/CD pipeline. A workflow that runs the tests and refuses to deploy when a metric drops is the whole of Unit 4 in about thirty lines.

  5. Understand what drift is not. Retraining on the drift in these notes gained 0.0016 accuracy, because the inputs moved and the relationship did not. Knowing when retraining will not help is the Unit 5 answer worth having.

  6. Read Unit 5's governance material properly. GDPR, CCPA and Responsible AI are examinable, they are the part with no code, and they are what makes this course different from Machine Learning.

If you read one thing

Unit 5's section on the three kinds of drift, and then run 12_serve_drift_govern.py.

Data drift, concept drift and label drift are routinely conflated, they need different detectors, and only one of them is detectable before the damage is done. The experiment demonstrates the distinction rather than asserting it — which is why the retraining loop it builds barely helps.

Units in this Course

UNIT 1

Foundations of Data Engineering

What data engineering is, and why 'systems' and 'maintenance' are the load-bearing words; the activities from ingestion to monitoring; the data lifecycle against the data ENGINEERING lifecycle, with the five undercurrents and why they are drawn underneath; the evolution of the role; ETL against ELT and why ELT won; technical against business responsibilities, internal against external; how data engineering relates to data science, in both directions; and the measured demonstration that a pipeline reporting a 1% error is more dangerous than one that crashes.

UNIT 2

Data Architecture and Distributed Systems

Enterprise, data and solution architecture; the principles of good architecture and why reversibility deserves your design effort; availability, reliability, RTO and RPO, with what each nine actually costs; tiers; monolith against microservices MEASURED, including the honest admission that microservices are slower and the one-database-per-service cost; event-driven architecture and batch against streaming ingestion measured at a 160x latency difference; dead-letter queues; hybrid cloud, multicloud and edge; and TCO, where the licence fee is rarely the largest line.

UNIT 3

MLOps Fundamentals

How MLOps differs from DevOps — code, data and model versioned together; training/serving skew and its architectural fix; data leakage and the three ways it happens; EDA, feature engineering and recording the base rate first; experiment tracking with real MLflow, and why the train/test gap belongs in the table; the four things that must be pinned for reproducibility, with the split's random_state demonstrated as the one people forget; model versioning; what DVC stores in git and what it does not; Responsible AI controls and what breaks at each scale.

UNIT 4

Model Deployment and CI/CD Pipelines

What production-ready means concretely, and why returning 400 with a reason matters more than the model; dev, staging and production, and putting the differences in configuration; CI/CD for ML including the two stages with no software equivalent — data validation and the metric gate; why a non-deterministic pipeline makes CI meaningless; batch, online, streaming and embedded deployment; canary, blue-green, A/B and shadow releases; the seven Docker traps; layer caching and image size; and Kubernetes readiness against liveness probes.

UNIT 5

Monitoring, Feedback Loops and Governance

Data, concept and label drift told apart, and the asymmetry that only one is detectable early; PSI and the KS test scored against drift injected at a known magnitude — 4 of 5 batches caught with no false alarms and a one-batch lag; why statistical significance is not operational significance; ground truth evaluation and the partial-label feedback trap; the seven-step retraining loop and the step that must stay human; metrics against logs, the four Prometheus types, and why latency needs a histogram; GDPR, CCPA, GxP and the EU AI Act; why Article 17 breaks trained models; fairness measured with a control condition; and a model risk management template.

PRACTICE

Practice

Exam-style questions with fully worked solutions.

LAB

Lab

Every prescribed lab experiment, with code and expected output.