Skip to the content
On this page
  1. The one thing to understand before anything else
  2. What runs here, and what does not
  3. Course objectives (verbatim)
  4. The five units
  5. Also here
  6. Cross-course connections
  7. Textbooks
  8. How to study this course
  9. Units in this Course

Part of the data-platform path: Big Data Technologies, Cloud Computing for Data Science, Time Series Analysis and Forecasting and Data Engineering and MLOps.


The one thing to understand before anything else

The cloud is not "someone else's computer". It is someone else's computer billed by the second, and that changes what is worth building.

Every technical decision in this course is downstream of one economic fact:

NOTE

Capacity you no longer need can be given back.

That is the whole thing. It is why virtualization matters (you cannot hand back half a physical server), why autoscaling exists, why serverless exists, why storage has six price tiers, and why "just leave it running" is the most expensive habit in the subject.

The old assumption What replaces it
Buy for the peak, own it for three years Rent the peak, hourly
Capacity is a capital decision Capacity is an API call
An idle server is sunk cost An idle server is a bill
Storage is a disk you bought Storage is six tiers with penalties
Security is a perimeter Security is identity

⚠️ The honest framing

This course names specific products, and the products change. SageMaker Studio replaced SageMaker Notebooks, Stackdriver became Cloud Operations, and the syllabus's product list will read as dated within a few years.

What does not change is underneath: service models, virtualization, storage classes, identity and policy evaluation, the batch/stream split, and the cost arithmetic. Learn those as permanent and the products as examples, and say so in the exam.


What runs here, and what does not

There is no cloud account for this repository, and none will be created. Signing up requires a payment card and accepts a billing relationship, which is not something a study repository should do on anyone's behalf.

So this course is the most explicitly split of the ten:

Runs for real Documented, NOT EXECUTED
IAM's policy evaluation algorithm, implemented and exercised AWS, Azure, GCP consoles
Object-store key semantics — prefixes, no directories, copy-plus-delete S3, Blob, Cloud Storage
All the pricing arithmetic — storage classes, egress, per-TB, per-node-hour the billing console
Hypervisor overcommit, and where it breaks VMware Workstation
A real web server serving a real page over TCP Apache on a cloud VM
A real ETL pipeline into a real columnar warehouse Glue, Redshift, BigQuery
An autoscaling control loop, measured honestly CloudWatch
A real model, and a real AutoML search SageMaker, Vertex, Azure ML
A real HTTP endpoint serving that model, called over the network a SageMaker endpoint

That is more than it sounds, because most of what this course teaches is not proprietary. IAM's three evaluation rules, the fact that an object store has no directories, and the arithmetic that decides between serverless and provisioned are all implementable — and all implemented, in labs/course-13b-cloud/.

Every file that cannot run says *** NOT EXECUTED *** at the top, names the service it needs, and points at the runnable half. tools/data-science/run_cloud_labs.py asserts the marker is still there.


Course objectives (verbatim)

  1. Introduce the fundamentals of cloud computing and its role in data science.
  2. Provide understanding of virtualization, service, and deployment models.
  3. Familiarize students with cloud storage, data management, and databases.
  4. Expose students to cloud-based big data and machine learning platforms.
  5. Train students in building, deploying, and monitoring ML pipelines on the cloud.

The five units

Unit Question it answers
1 What is the cloud, and what are you actually renting?
2 How is one machine turned into many, and whose machine is it?
3 Where does the data live, and what does each choice cost?
4 What does a managed ML platform actually give you?
5 How do you train, deploy and keep it working?

Unit 3 is the load-bearing one. Storage decisions are where cloud bills are made and lost, and they are the most examinable arithmetic in the course.


Also here

Cross-course connections

From To What is shared
Business Intelligence Tools (BI) Unit 3, experiment 12 The same nine-row star schema, imported not copied. ₹10,360 for South is now produced by four engines.
Big Data Technologies Units 3 and 4 Column projection and partition pruning saved time there; here they save money, on the same mechanism.
Machine Learning (ML) Units 4 and 5 Identical scikit-learn. The cloud changes the packaging, not the algorithm — and the base-rate argument survives intact.
Document Oriented Database (MongoDB) Unit 3 Cosmos DB's consistency levels are CAP as a dropdown, with a price per level.
Database Management Systems (DBMS) Unit 3 A cloud warehouse is columnar, indexless, and billed per byte scanned. Every row of that comparison is a departure from Database Management Systems.

Textbooks

The syllabus gives a single combined Text / Reference list:

WATCH OUT

⚠️ Item 4 of the list is empty

The prescribed list runs 1, 2, 3, 4, 5 with nothing beside the 4 — a title has been lost, and the four books above are items 1, 2, 3 and 5. See review finding D22.

None of the four is free. AWS, Azure and Google Cloud all publish their own documentation and free tiers, and for Units 4 and 5 the vendor documentation is more current than any of these books.

How to study this course

  1. Learn the three service models by what YOU manage, not by examples. The examples change; the boundary does not.

  2. Do the cost arithmetic by hand. 1 TB egress, 1 TB in each storage class, a serverless-vs-provisioned break-even. These are the calculations that get examined.

  3. Learn IAM's three rules and be able to apply them. Explicit deny wins; otherwise allow; otherwise deny. Almost every access question is these three.

  4. Learn one comparison table per unit. IaaS/PaaS/SaaS; the four deployment models; block/file/object; real-time/serverless/batch.

  5. Run the labs. Seven programs, including a real web server, a real ETL into a real warehouse, and a real model served over a real HTTP endpoint.

KEY INSIGHT

The two sentences that carry the course

Capacity you no longer need can be given back. That is why the cloud exists.

Nothing you forget to switch off will switch itself off. That is why cloud bills surprise people, and it is worth a mark in almost any question about cost.

Units in this Course

UNIT 1

Introduction to Cloud Computing

The NIST definition and the sentence underneath it; the evolution from time-sharing through grid and utility computing; the five essential characteristics, and how to use them as a test; SOA, web services and why REST replaced SOAP; the four-part architecture and the control plane; IaaS, PaaS and SaaS by what the customer manages; continuous delivery, blue/green and canary.

UNIT 2

Virtualization and Deployment Models

Why virtualization is what makes the cloud possible; partitioning, isolation and encapsulation; type 1 against type 2 hypervisors, and containers; the six types of virtualization; overcommitment measured, and why CPU degrades gracefully while memory falls off a cliff; public, private, community and hybrid; the role of the cloud in data science, and what it does not fix.

UNIT 3

Cloud Storage and Data Management

Block, file and object storage compared on access unit, sharing and cost; why an object store has no directories and no rename; storage classes, minimum durations and the retrieval fees that reverse the discount; egress and data gravity; backup, archiving, DR and content delivery; key-value databases and their limitations; batch against streaming; cloud data warehouses, bytes scanned, and the break-even.

UNIT 4

Cloud Platforms for Data Science and ML

What the cloud changes about machine learning and what it does not; the benefits and the catch attached to each; AIaaS and GPUaaS as SaaS and IaaS for particular things; managed platforms compared — SageMaker, Azure ML, Vertex AI; the model registry, feature store and experiment tracking; AutoML run for real, what it costs, and the seven things it cannot do.

UNIT 5

Training and Deployment of ML on the Cloud

Choosing a platform — pipeline support, scale-up against scale-out, framework support, pre-tuned services; the six steps and the failure mode at each; the container contract and why /ping must not run the model; real-time, serverless and batch inference with cost figures; monitoring, alarming on the tail, and autoscaling measured honestly; drift, retraining and case studies.

PRACTICE

Practice

Exam-style questions with fully worked solutions.

LAB

Lab

Every prescribed lab experiment, with code and expected output.

Next course in learning order: Data Mining Machine learning & AI