15 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.
Code lives in labs/course-13b-cloud/.
NOTE
There is no cloud account for this repository, and none will be created.
Signing up for AWS, Azure or GCP requires a payment card and accepts a billing relationship. That is not a thing a study repository should do on anyone's behalf, so no provider was ever contacted and no claim in these notes about a provider's behaviour was demonstrated here.
| Half | Files | Status |
|---|---|---|
| The console and CLI steps | 14 Markdown files | NOT EXECUTED, at the top of every one: each experiment's 4. Programme is its procedure, and 5. Execution and Results says so in a box |
| The verification | 7 programs | Executed and asserted by tools/data-science/run_cloud_labs.py; what each printed is under the experiment it is named for |
pip install -r tools/requirements.txt
python3 tools/data-science/run_cloud_labs.py
AND YET A SURPRISING AMOUNT REALLY RUNS
Most of what this course teaches is not proprietary:
| Runs for real | What it is |
|---|---|
| IAM policy evaluation | the actual algorithm, in iam.py |
| Object-store semantics | prefixes, no directories, copy-plus-delete, versioning |
| All the pricing arithmetic | storage classes, egress, per-TB, per-node-hour |
| Hypervisor overcommit | and the point at which it fails |
| A real web server | serving a real page over TCP, fetched back |
| A real ETL pipeline | SQLite → transform → DuckDB, with an audit trail |
| An autoscaling control loop | measured, including where it loses |
| A real model and a real AutoML search | scikit-learn, 25 real fits |
| A real REST endpoint | serving that model, called over the network |
Nothing is claimed that was not executed. Every .md file names the
service it needs and the runnable half that verifies its logic, and the runner
asserts the marker is still present.
Three of the programs time something — a training job, a model search, an endpoint's latency.
A timing measures the machine at a moment, so those lines differ from run to run, and
capture_lab_outputs.py --check sets them aside; every other line must repeat exactly.
Experiments 8, 9 and 12 use Business Intelligence Tools' star schema, imported not copied.
₹10,360 for South is now produced by Business Intelligence Tools' DAX, Big Data Technologies' Hive,
Big Data Technologies' Spark and this course's DuckDB — four engines, nine facts,
and verify_all.sh fails if any of them drifts.
Create a virtual machine in VMware Workstation.
Run the new-VM wizard, make the three choices that matter, and see where memory overcommit fails.
The procedure, on the console and CLI, 01_create_vm.md:
The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:
THE THREE WIZARD CHOICES THAT GET PEOPLE
Memory. A type 2 hypervisor does not balloon aggressively, so the overcommit that works in a datacentre does not work on a laptop. Give a 16 GB laptop's VM 12 GB and the whole machine swaps.
Disk: pre-allocate or grow. Pre-allocating writes 40 GB immediately and is faster after; growing on demand is what you want on a laptop.
NAT, Bridged or Host-only. NAT means the LAN cannot reach the guest — and that is experiment 2's most common failure.
The procedure, on the console and CLI, 01_create_vm.md:
# Experiment 1 -- create a virtual machine in VMware Workstation
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`01_vm_and_hosting.py`, which models overcommit and measures where it breaks**.
---
<!-- Step 1: Run the wizard -->
## The wizard, and what each choice actually means
| Step | Choice | Why it matters |
|---|---|---|
| New Virtual Machine | **Custom (advanced)** | Typical hides the disk and network options you need |
| Hardware compatibility | Workstation 17.x | older only if you must move the VM to an older host |
| Guest OS install | **Install later** | attach the ISO after, so the wizard does not run an unattended install |
| Guest OS | Linux → Ubuntu 64-bit | picks sensible defaults for the virtual chipset |
| Processors | 2 cores | more than the host has *physical* cores makes it slower, not faster |
| Memory | 4096 MB | see the overcommit note below |
| Network | **NAT** | the guest shares the host's IP; Bridged gives it one on your LAN |
| Disk controller | NVMe (or SCSI) | IDE is slow and there is no reason for it |
| Disk | 40 GB, **split**, **not** pre-allocated | see the disk note below |
<!-- Step 2: Make the three choices that matter -->
## The three choices that get people
**1. Memory.** The host must keep enough for itself. Giving a 16 GB laptop's
VM 12 GB leaves 4 GB for Windows, the hypervisor and your browser, and the
whole machine swaps. **Type 2 hypervisors do not balloon aggressively**, so
the overcommit that works in a datacentre does not work here.
**2. Disk: "Allocate all disk space now" vs "split into multiple files".**
Pre-allocating writes 40 GB immediately and is faster afterwards. Not
pre-allocating grows on demand and is what you want on a laptop. Splitting
into 2 GB files matters only for filesystems that cannot hold a 40 GB file
(FAT32) — and for copying the VM to a USB stick.
**3. NAT vs Bridged vs Host-only.**
| Mode | The guest gets | Reachable from the LAN? |
|---|---|---|
| **NAT** | a private IP behind the host | **no**, unless you forward a port |
| **Bridged** | an IP from your router's DHCP | **yes** |
| Host-only | a private IP, no internet | no |
**Choose NAT and then wonder why nobody can reach your web server** — that is
experiment 2's most common failure, and the fix is either Bridged or a port
forward in `Edit → Virtual Network Editor`.
<!-- Step 3: Finish the install -->
## After the install
```bash
sudo apt update && sudo apt install -y open-vm-tools open-vm-tools-desktop
# ^ shared clipboard, drag-and-drop, correct screen resolution
ip addr show # note the guest's IP -- you need it for experiment 2
free -h ; nproc ; df -h
```
**Take a snapshot before you install anything else.** `VM → Snapshot → Take
Snapshot`. A snapshot is the only reason experimenting in a VM is safe, and
it is the feature that has no cheap equivalent on physical hardware.
## The link forward
Every EC2 instance, Azure VM and GCE instance is a guest on a **type 1**
hypervisor — ESXi, Hyper-V, KVM or AWS's Nitro. **The cloud is this
experiment at rack scale**, with the wizard replaced by an API call. Knowing
what the wizard was choosing is what makes an instance type comprehensible.
The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:
"""Experiments 1, 2 and 7 -- a virtual machine, a web server on it, and a
notebook environment.
VMware Workstation is not installed here, so `01_create_vm.md` carries the
wizard steps, marked as not run here. But the two things those experiments actually
teach DO run:
* virtualization is RESOURCE MULTIPLEXING, and the interesting behaviour is
overcommit -- modelled and measured below.
* experiment 2 hosts a page on a server. THIS SCRIPT REALLY DOES THAT,
with Python's own HTTP server standing in for Apache, and fetches the
page back over TCP to prove it.
* experiment 7 runs a notebook. Papermill and Jupyter are not installed,
so the script executes the same cells directly and asserts the outputs,
which is what a notebook test does anyway.
"""
import http.server
import json
import os
import socketserver
import tempfile
import threading
import urllib.error
import urllib.request
import fixtures as f
HOST = "127.0.0.1"
PORT = 0 # let the OS pick a free port -- see experiment_2
# ------------------------------------------------------------ experiment 1
def allocate(host_ram_gb, host_vcpu, vms, ballooning=True):
"""Place VMs on a host and report what the hypervisor actually does.
Two facts drive everything:
* vCPUs are TIME-SLICED, so you can allocate far more than you have.
* RAM is not, at least not for free -- overcommitted memory is backed
by ballooning, page sharing and finally SWAP, which is a cliff.
"""
ram_alloc = sum(v["ram"] for v in vms)
cpu_alloc = sum(v["vcpu"] for v in vms)
ram_used = sum(v["ram"] * v["active"] for v in vms)
reclaimed = (ram_alloc - ram_used) if ballooning else 0
pressure = max(0.0, ram_used - host_ram_gb)
return {
"ram_allocated": ram_alloc,
"ram_ratio": ram_alloc / host_ram_gb,
"cpu_ratio": cpu_alloc / host_vcpu,
"ram_actually_touched": ram_used,
"reclaimed_by_ballooning": reclaimed,
"swapping_gb": pressure,
}
def experiment_1():
print("\n --- experiment 1: the virtual machine")
print(f" {'':<22}{'type 1 (bare metal)':<26}{'type 2 (hosted)'}")
for label, t1, t2 in (
("runs on", "the hardware directly", "on top of an OS"),
("examples", "ESXi, Hyper-V, KVM, Xen", "VMware Workstation, VirtualBox"),
("overhead", "a few percent", "noticeably more"),
("used for", "datacentres, THE CLOUD", "a laptop, this experiment"),
("boots", "instead of an OS", "as an application")):
print(f" {label:<22}{t1:<26}{t2}")
print(""" every EC2 instance, every Azure VM and every GCE instance
is a guest on a TYPE 1 hypervisor. The whole cloud is this
experiment, at rack scale -- which is why it is experiment 1""")
host_ram, host_cpu = 32, 8
vms = [
{"name": "web-1", "ram": 8, "vcpu": 4, "active": 0.35},
{"name": "web-2", "ram": 8, "vcpu": 4, "active": 0.30},
{"name": "db-1", "ram": 16, "vcpu": 4, "active": 0.90},
{"name": "batch", "ram": 16, "vcpu": 8, "active": 0.20},
]
print(f"\n a {host_ram} GB / {host_cpu} vCPU host, four guests:")
print(f" {'vm':<10}{'RAM':>6}{'vCPU':>6}{'active':>9}")
for v in vms:
print(f" {v['name']:<10}{v['ram']:>5} G{v['vcpu']:>6}"
f"{v['active']:>8.0%}")
r = allocate(host_ram, host_cpu, vms)
print(f"\n allocated RAM : {r['ram_allocated']} GB on a {host_ram} GB host "
f"({r['ram_ratio']:.2f}x)")
print(f" allocated vCPU : {sum(v['vcpu'] for v in vms)} on {host_cpu} "
f"({r['cpu_ratio']:.2f}x)")
print(f" RAM actually touched : {r['ram_actually_touched']:.1f} GB")
print(f" reclaimed by ballooning : {r['reclaimed_by_ballooning']:.1f} GB")
print(f" swapping : {r['swapping_gb']:.1f} GB")
assert r["ram_ratio"] > 1 and r["cpu_ratio"] > 1
assert r["swapping_gb"] == 0
print(f""" {r['ram_allocated']} GB allocated on a {host_ram} GB host and
{sum(v['vcpu'] for v in vms)} vCPUs on {host_cpu}, and nothing is swapping -- because
the guests only TOUCH {r['ram_actually_touched']:.1f} GB.
Overcommit works on the same bet an airline makes, and it is
why a cloud provider can sell more capacity than it owns""")
print("\n now the batch job wakes up (20% -> 95% active):")
busy = [dict(v, active=0.95 if v["name"] == "batch" else v["active"])
for v in vms]
r2 = allocate(host_ram, host_cpu, busy)
print(f" RAM actually touched : {r2['ram_actually_touched']:.1f} GB")
print(f" swapping : {r2['swapping_gb']:.1f} GB")
assert r2["swapping_gb"] > 0
print(f""" {r2['swapping_gb']:.1f} GB OVER, AND NOW EVERY GUEST IS SLOW -- not just
the batch job. Memory overcommit fails as a CLIFF, and it
fails for the neighbours: this is the 'noisy neighbour'
problem, and it is why cloud instance types quote DEDICATED
memory and only burstable CPU.
CPU overcommit degrades gracefully because time-slicing
shares; RAM does not, because a page is either resident or
it is not""")
# ------------------------------------------------------------ experiment 2
PAGE = """<!doctype html>
<title>Sales dashboard</title>
<h1>Retail sales</h1>
<table>
<tr><th>Region</th><th>Revenue</th></tr>
{rows}
</table>
<p>Total: {total}</p>
"""
def experiment_2():
print("\n --- experiment 2: host a page on the server (this RUNS)")
doc_root = tempfile.mkdtemp(prefix="cloud13b_www_")
by_region = (f.SALES_DF.groupby("region")["revenue"].sum()
.sort_values(ascending=False))
rows = "\n".join(f"<tr><td>{k}</td><td>{v:,.0f}</td></tr>"
for k, v in by_region.items())
html = PAGE.format(rows=rows, total=f"{f.total_revenue():,.0f}")
with open(os.path.join(doc_root, "index.html"), "w") as fh:
fh.write(html)
with open(os.path.join(doc_root, "data.json"), "w") as fh:
json.dump({k: float(v) for k, v in by_region.items()}, fh)
class Quiet(http.server.SimpleHTTPRequestHandler):
def __init__(self, *a, **kw):
super().__init__(*a, directory=doc_root, **kw)
def log_message(self, *a):
pass
class Reusable(socketserver.TCPServer):
allow_reuse_address = True # or a re-run hits TIME_WAIT
with Reusable((HOST, PORT), Quiet) as httpd:
port = httpd.server_address[1]
t = threading.Thread(target=httpd.serve_forever, daemon=True)
t.start()
# [Changed: these printed the temporary folder's name and the port, which
# the system picks afresh on every run.]
print(" document root : a new temporary folder, cloud13b_www_...")
print(f" serving : http://{HOST}, on a port the system chose (a REAL server)")
with urllib.request.urlopen(f"http://{HOST}:{port}/") as resp:
served = resp.read().decode()
ctype = resp.headers["Content-Type"]
status = resp.status
print(f" GET / -> {status}, {ctype}, "
f"{len(served)} bytes")
assert status == 200 and "text/html" in ctype
assert "Retail sales" in served and "10,360" in served
with urllib.request.urlopen(f"http://{HOST}:{port}/data.json") as resp:
data = json.loads(resp.read())
jtype = resp.headers["Content-Type"]
print(f" GET /data.json-> 200, {jtype}, {data}")
assert data["South"] == 10360.0 and jtype == "application/json"
try:
urllib.request.urlopen(f"http://{HOST}:{port}/missing.html")
raise AssertionError("should have 404ed")
except urllib.error.HTTPError as exc:
print(f" GET /missing -> {exc.code}")
assert exc.code == 404
httpd.shutdown()
os.remove(os.path.join(doc_root, "index.html"))
os.remove(os.path.join(doc_root, "data.json"))
os.rmdir(doc_root)
print(""" a page was written to a document root, served over TCP,
fetched back, and its CONTENT-TYPE checked. That is the whole
of experiment 2; Apache under XAMPP adds virtual hosts,
.htaccess, PHP and TLS, and the shape is identical.
Note the Content-Type header. A browser renders index.html
because the server SAID text/html -- get that wrong and the
browser downloads your page instead of showing it, which is
the commonest 'my site is broken' on a fresh VM""")
print(f"\n {'concern':<24}{'on your VM':<26}{'managed (S3/App Service)'}")
for c, vm, mg in (
("who patches Apache", "YOU, monthly", "the provider"),
("TLS certificate", "certbot, renewals", "issued and rotated"),
("scaling", "a bigger VM", "automatic"),
("a static site costs", "a VM, hourly", "cents per GB stored"),
("you control", "everything", "very little")):
print(f" {c:<24}{vm:<26}{mg}")
print(""" a STATIC site on a VM is the clearest case of paying for
a general-purpose computer to do something an object store
does for cents. Hosting index.html on S3 + CloudFront costs
less than the VM's first hour""")
# ------------------------------------------------------------ experiment 7
def experiment_7():
print("\n --- experiment 7: the notebook environment")
cells = [
("import pandas as pd; import fixtures as f",
lambda ns: ns.update({"df": f.SALES_DF}) or "ok"),
("df.shape", lambda ns: ns["df"].shape),
("df.groupby('region')['revenue'].sum().to_dict()",
lambda ns: {k: float(v) for k, v in
ns["df"].groupby("region")["revenue"].sum().items()}),
("df['revenue'].sum()", lambda ns: float(ns["df"]["revenue"].sum())),
]
ns, outputs = {}, []
for src, fn in cells:
out = fn(ns)
outputs.append(out)
shown = str(out)
print(f" In [{len(outputs)}]: {src}")
print(f" Out [{len(outputs)}]: "
f"{shown[:60]}{'...' if len(shown) > 60 else ''}")
assert outputs[1] == (9, 19)
assert outputs[2]["South"] == 10360.0
assert outputs[3] == f.total_revenue()
print(""" four cells, executed in order, every output asserted.
That is what a notebook TEST looks like -- papermill or
nbconvert --execute do exactly this in CI, and a notebook
nobody executes in CI is a notebook that has already
drifted""")
print(f"\n {'':<24}{'Colab':<24}{'notebook on a cloud VM'}")
for label, colab, vm in (
("costs", "free tier, then paid", "the INSTANCE, hourly"),
("data access", "upload, or mount Drive", "IAM role, no keys"),
("stops when", "idle ~90 min", "NEVER -- you stop it"),
("state on stop", "LOST", "kept on the EBS volume"),
("GPU", "when available", "the one you pay for"),
("private data", "a policy question", "inside your VPC")):
print(f" {label:<24}{colab:<24}{vm}")
idle = f.EC2["m5.xlarge"] * f.HOURS_PER_MONTH
print(f"\n an m5.xlarge notebook left running: ${idle:,.2f}/month")
assert idle > 100
print(""" 'STOPS WHEN: NEVER' is the row that costs money. Colab
disconnecting is an annoyance; a cloud notebook not
disconnecting is a bill. Set an idle-shutdown lifecycle
policy on day one -- SageMaker supports one, and it is the
single most useful thing you can configure""")
def main():
print(" Experiments 1, 2 and 7 -- VM, web server and notebook")
# Step 1: Experiment 1: allocate the guests, and overcommit
experiment_1()
# Step 2: Experiment 2: serve a page
experiment_2()
# Step 3: Experiment 7: run the notebook's cells
experiment_7()
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 01_create_vm.md:
NOT RUN HERE
01_create_vm.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:
OUTPUT
Experiments 1, 2 and 7 -- VM, web server and notebook
--- experiment 1: the virtual machine
type 1 (bare metal) type 2 (hosted)
runs on the hardware directly on top of an OS
examples ESXi, Hyper-V, KVM, Xen VMware Workstation, VirtualBox
overhead a few percent noticeably more
used for datacentres, THE CLOUD a laptop, this experiment
boots instead of an OS as an application
every EC2 instance, every Azure VM and every GCE instance
is a guest on a TYPE 1 hypervisor. The whole cloud is this
experiment, at rack scale -- which is why it is experiment 1
a 32 GB / 8 vCPU host, four guests:
vm RAM vCPU active
web-1 8 G 4 35%
web-2 8 G 4 30%
db-1 16 G 4 90%
batch 16 G 8 20%
allocated RAM : 48 GB on a 32 GB host (1.50x)
allocated vCPU : 20 on 8 (2.50x)
RAM actually touched : 22.8 GB
reclaimed by ballooning : 25.2 GB
swapping : 0.0 GB
48 GB allocated on a 32 GB host and
20 vCPUs on 8, and nothing is swapping -- because
the guests only TOUCH 22.8 GB.
Overcommit works on the same bet an airline makes, and it is
why a cloud provider can sell more capacity than it owns
now the batch job wakes up (20% -> 95% active):
RAM actually touched : 34.8 GB
swapping : 2.8 GB
2.8 GB OVER, AND NOW EVERY GUEST IS SLOW -- not just
the batch job. Memory overcommit fails as a CLIFF, and it
fails for the neighbours: this is the 'noisy neighbour'
problem, and it is why cloud instance types quote DEDICATED
memory and only burstable CPU.
CPU overcommit degrades gracefully because time-slicing
shares; RAM does not, because a page is either resident or
it is not
--- experiment 2: host a page on the server (this RUNS)
document root : a new temporary folder, cloud13b_www_...
serving : http://127.0.0.1, on a port the system chose (a REAL server)
GET / -> 200, text/html, 225 bytes
GET /data.json-> 200, application/json, {'South': 10360.0, 'North': 2520.0}
GET /missing -> 404
a page was written to a document root, served over TCP,
fetched back, and its CONTENT-TYPE checked. That is the whole
of experiment 2; Apache under XAMPP adds virtual hosts,
.htaccess, PHP and TLS, and the shape is identical.
Note the Content-Type header. A browser renders index.html
because the server SAID text/html -- get that wrong and the
browser downloads your page instead of showing it, which is
the commonest 'my site is broken' on a fresh VM
concern on your VM managed (S3/App Service)
who patches Apache YOU, monthly the provider
TLS certificate certbot, renewals issued and rotated
scaling a bigger VM automatic
a static site costs a VM, hourly cents per GB stored
you control everything very little
a STATIC site on a VM is the clearest case of paying for
a general-purpose computer to do something an object store
does for cents. Hosting index.html on S3 + CloudFront costs
less than the VM's first hour
--- experiment 7: the notebook environment
In [1]: import pandas as pd; import fixtures as f
Out [1]: ok
In [2]: df.shape
Out [2]: (9, 19)
In [3]: df.groupby('region')['revenue'].sum().to_dict()
Out [3]: {'North': 2520.0, 'South': 10360.0}
In [4]: df['revenue'].sum()
Out [4]: 12880.0
four cells, executed in order, every output asserted.
That is what a notebook TEST looks like -- papermill or
nbconvert --execute do exactly this in CI, and a notebook
nobody executes in CI is a notebook that has already
drifted
Colab notebook on a cloud VM
costs free tier, then paid the INSTANCE, hourly
data access upload, or mount Drive IAM role, no keys
stops when idle ~90 min NEVER -- you stop it
state on stop LOST kept on the EBS volume
GPU when available the one you pay for
private data a policy question inside your VPC
an m5.xlarge notebook left running: $140.16/month
'STOPS WHEN: NEVER' is the row that costs money. Colab
disconnecting is an annoyance; a cloud notebook not
disconnecting is a bill. Set an idle-shutdown lifecycle
policy on day one -- SageMaker supports one, and it is the
single most useful thing you can configure
Overcommit, measured. A 32 GB / 8 vCPU host with four guests:
allocated RAM : 48 GB on a 32 GB host (1.50x)
allocated vCPU : 20 on 8 (2.50x)
RAM actually touched : 22.8 GB
reclaimed by ballooning : 25.2 GB
swapping : 0.0 GB
48 GB allocated on 32 GB, and nothing is swapping, because the guests only touch 22.8 GB. Overcommit works on the same bet an airline makes.
Then the batch job wakes up (20% → 95% active):
RAM actually touched : 34.8 GB
swapping : 2.8 GB
Now every guest is slow — not just the batch job.
THE ASYMMETRY TO REMEMBER
NOTE
CPU overcommit degrades gracefully. Memory overcommit fails as a cliff.
CPU is time-sliced, so twice the demand means half the speed for everyone. A memory page is either resident or on disk, and the difference is a factor of thousands. That is the "noisy neighbour" problem, and it is why cloud instance types quote dedicated memory and only burstable CPU.
Changed: the program printed the temporary document root's full name and the web server's port, which the system picks afresh on every run; it now says what they are.
RESULT
48 GB allocated on a 32 GB host swaps nothing while the guests touch 22.8 GB; when the batch job wakes, 2.8 GB swaps and every guest slows.
Install and configure Apache (or XAMPP) on the VM, and host a page.
Serve a page, check the header that decides whether a browser shows it, and compare a VM with an object store.
The procedure, on the console and CLI, 02_web_server.md:
WHAT RUNS
The page is really served, in 01_vm_and_hosting.py's experiment 2:
GET / -> 200, text/html, 225 bytes
GET /data.json -> 200, application/json, {'South': 10360.0, 'North': 2520.0}
GET /missing -> 404
A page was written to a document root, served over TCP, fetched back, and
its Content-Type checked. That is the whole of experiment 2; Apache under
XAMPP adds virtual hosts, .htaccess, PHP and TLS, and the shape is identical.
The procedure, on the console and CLI, 02_web_server.md:
# Experiment 2 -- install and configure Apache/XAMPP on the VM and host a page
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`01_vm_and_hosting.py`, which really does serve a page over TCP and fetch it back**.
---
<!-- Step 1: Install Apache, on Linux -->
## Linux (Apache directly)
```bash
sudo apt update && sudo apt install -y apache2
sudo systemctl enable --now apache2
systemctl status apache2
sudo tee /var/www/html/index.html >/dev/null <<'EOF'
<!doctype html>
<title>Sales dashboard</title>
<h1>Retail sales</h1>
EOF
curl -I http://localhost/ # 200, and check the Content-Type
```
From the HOST machine, using the guest's IP from experiment 1:
```
http://192.168.x.x/
```
**If that times out**, work through in this order:
1. `sudo ufw allow 80/tcp` — the guest firewall
2. the network mode: **NAT means the LAN cannot reach the guest** (see
experiment 1). Switch to Bridged, or forward a port.
3. `sudo ss -tlnp | grep :80` — is Apache bound to `0.0.0.0` or only to
`127.0.0.1`?
<!-- Step 2: Or install XAMPP, on Windows -->
## Windows (XAMPP/WAMP)
Install XAMPP, start **Apache** from the control panel, put the page in
`C:\xampp\htdocs\`, browse to `http://localhost/`.
**Apache will not start** almost always because **port 80 is taken** — by IIS,
by Skype (historically), or by another web server. Change
`Listen 80` to `Listen 8080` in `httpd.conf` and browse to
`http://localhost:8080/`.
<!-- Step 3: Configure it -->
## The configuration worth understanding
| Directive | Where | What it does |
|---|---|---|
| `DocumentRoot` | `apache2.conf` / `httpd.conf` | which directory is served |
| `Listen` | `ports.conf` | which port |
| `<VirtualHost>` | `sites-available/` | several sites on one IP, by hostname |
| `DirectoryIndex` | `apache2.conf` | which file `/` serves |
| `AddType` | `mime.types` | the **Content-Type** header |
**`AddType` is the one that bites.** A browser renders your page because the
server *said* `text/html`. Serve it as `text/plain` and the browser shows
source; serve it as `application/octet-stream` and the browser downloads it.
The runnable half asserts this header for exactly that reason.
<!-- Step 4: Enable TLS -->
## Enabling TLS
```bash
sudo apt install -y certbot python3-certbot-apache
sudo certbot --apache -d example.com # needs a real domain pointing here
```
**You cannot get a certificate for `192.168.1.50`.** Certificate authorities
issue for names they can validate, so a VM on your LAN gets a self-signed
certificate and a browser warning — which is the correct behaviour, not a bug.
<!-- Step 5: Compare it with an object store -->
## The comparison that ends the experiment
| | This VM | S3 + CloudFront |
|---|---|---|
| Patch Apache | **you, monthly** | not your problem |
| TLS certificate | certbot, and renewals | issued and rotated |
| Survives your laptop closing | **no** | yes |
| Handles a traffic spike | no | yes |
| Cost for a static page | a VM, hourly | **cents per GB** |
**A static site on a VM is a general-purpose computer doing an object store's
job.** That is the point the experiment makes by having you do it the hard way
once.
The procedure, on the console and CLI, 02_web_server.md:
NOT RUN HERE
02_web_server.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
THE HEADER THAT DECIDES WHETHER YOUR SITE WORKS
A browser renders index.html because the server said text/html. Get
AddType wrong and the browser downloads your page instead of showing it —
the commonest "my site is broken" on a fresh VM, and the reason the runnable
half asserts the header rather than just the body.
And the comparison that ends the experiment:
| This VM | S3 + CloudFront | |
|---|---|---|
| Patch Apache | you, monthly | not your problem |
| TLS certificate | certbot, and renewals | issued and rotated |
| Survives your laptop closing | no | yes |
| A static page costs | a VM, hourly | cents per GB |
A static site on a VM is a general-purpose computer doing an object store's job, and the experiment makes the point by having you do it the hard way once.
The Python model for this experiment, 01_vm_and_hosting.py, is shown in full under Experiment 1, with what it printed.
RESULT
A page written to a document root was served over TCP and fetched back as text/html; a missing page gave 404.
Create and configure a cloud account on the AWS, Azure or GCP free tier.
Set the account up safely, and see exactly how IAM decides each request.
The procedure, on the console and CLI, 03_account_setup.md:
The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:
THE THREE RULES, AND THEY ARE THE WHOLE SUBJECT
Evaluated against a realistic policy set:
| Action | Resource | Result | Why |
|---|---|---|---|
s3:GetObject |
retail-lake/raw/sales.csv |
Allow | DataScientistRead |
s3:PutObject |
retail-lake/raw/sales.csv |
Deny | explicit deny in ProtectRawZone |
s3:PutObject |
retail-lake/models/model.pkl |
Allow | SageMakerExecution |
s3:GetObject |
other-bucket/secret.csv |
Deny | implicit — nothing matched |
Read rows 2 and 3 together. The same action on the same bucket is denied
under raw/ and allowed under models/, because a Deny scoped to one prefix
beats an Allow scoped to the bucket. That is how a data lake keeps a raw
zone immutable while the rest stays writable.
The procedure, on the console and CLI, 03_account_setup.md:
# Experiment 3 -- create and configure a cloud account (AWS/Azure/GCP free tier)
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`03_iam_and_account.py`, which implements and exercises IAM's evaluation algorithm**.
---
<!-- Step 1: Do the six things first -->
## Do these six things in this order, before anything else
1. **Enable MFA on the root account.** Then never use the root account again.
2. **Create an IAM user for yourself** with `AdministratorAccess`, and use
that.
3. **Set a budget alarm.** Billing → Budgets → a monthly cost budget with an
alert at 50%, 80% and 100%. **Do this on day one**, not after the bill.
4. **Enable Cost Explorer.** It takes 24 hours to populate, so turn it on now.
5. **Pick one region and stay in it.** Resources in another region are
invisible in the console and still bill.
6. **Install and configure the CLI:**
```bash
aws configure # access key, secret, default region, output format
aws sts get-caller-identity # who am I, really
aws configure list-profiles
```
<!-- Step 2: Use an IAM user, not root -->
## Root against IAM user
| | Root | IAM user |
|---|---|---|
| Created by | signing up | you |
| Can | **everything, including closing the account** | what its policies allow |
| MFA | **mandatory, immediately** | strongly recommended |
| Daily use | **never** | yes |
| Recovers | account-level problems only | — |
<!-- Step 3: Stay within the free tier -->
## The free tier, honestly
| Type | Duration | Examples |
|---|---|---|
| **12 months free** | from signup | 750 h/month t3.micro, 5 GB S3, 750 h RDS |
| **Always free** | forever | 1 M Lambda requests/month, 25 GB DynamoDB |
| **Trials** | short, one-off | SageMaker 250 h for 2 months |
**Three things bill you anyway, inside the free tier:**
- **Egress.** Ingress is free; downloading is not.
- **Anything larger than the free size.** A `t3.small` is not a `t3.micro`.
- **Resources you forgot.** An idle NAT Gateway is about $32/month, an
unattached Elastic IP is about $3.60/month, and a SageMaker endpoint is
about $70/month.
<!-- Step 4: Know the equivalents -->
## The equivalents, when the exam asks
| Concept | AWS | Azure | GCP |
|---|---|---|---|
| Object storage | S3 | Blob Storage | Cloud Storage |
| Block storage | EBS | Managed Disks | Persistent Disk |
| File storage | EFS | Azure Files | Filestore |
| VM | EC2 | Virtual Machines | Compute Engine |
| Managed SQL | RDS | SQL Database | Cloud SQL |
| Warehouse | Redshift | Synapse | **BigQuery** |
| Serverless functions | Lambda | Functions | Cloud Functions |
| Managed ML | SageMaker | Azure ML | Vertex AI |
| Monitoring | CloudWatch | Monitor | Cloud Monitoring |
| Identity | IAM | Entra ID | IAM |
**Learn the row, not the column.** Every provider has all of these, and an
exam answer that names the concept and gives one vendor's term for it is
worth more than a memorised AWS product list.
<!-- Step 5: Clean up -->
## Then clean up
```bash
aws ec2 describe-instances --query \
'Reservations[].Instances[?State.Name==`running`].[InstanceId,InstanceType]'
aws s3 ls
aws sagemaker list-endpoints
```
**Run those three at the end of every lab session.** Nothing else in this
course will save you as much money.
The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:
"""Experiments 3 and 10 -- cloud account setup, and IAM roles for SageMaker.
There is no cloud account for this repository and none will be created, so
`03_account_setup.md` and `10_sagemaker_notebook.md` carry the console
click-paths, marked as not run here.
What runs here is the part that is actually examinable: AWS's policy
evaluation algorithm, implemented in iam.py and exercised against a realistic
policy set. The three rules are the whole subject.
"""
import fixtures as f
from iam import evaluate
def show(policies, cases, title):
print(f"\n {title}")
print(f" {'action':<32}{'resource':<47}{'result':<7}why")
for action, resource in cases:
trace = {}
decision = evaluate(policies, action, resource, trace)
print(f" {action:<32}{resource:<47}"
f"{decision:<7}{trace['reason']}")
return {(a, r): evaluate(policies, a, r) for a, r in cases}
def main():
print(" Experiments 3 and 10 -- accounts, roles and IAM evaluation")
print("""
the three rules, in order:
1. an EXPLICIT DENY anywhere wins -- always, unconditionally
2. otherwise an ALLOW that matches grants access
3. otherwise DENY -- the IMPLICIT DENY""")
# Step 1: Evaluate the attached policies
cases = [
("s3:GetObject", "arn:aws:s3:::retail-lake/raw/sales.csv"),
("s3:PutObject", "arn:aws:s3:::retail-lake/raw/sales.csv"),
("s3:PutObject", "arn:aws:s3:::retail-lake/models/model.pkl"),
("s3:DeleteObject", "arn:aws:s3:::retail-lake/raw/sales.csv"),
("s3:GetObject", "arn:aws:s3:::other-bucket/secret.csv"),
("sagemaker:CreateTrainingJob", "arn:aws:sagemaker:*:*:training-job/x"),
]
got = show(f.POLICIES, cases, "the attached policies, evaluated:")
assert got[("s3:GetObject", "arn:aws:s3:::retail-lake/raw/sales.csv")] == "Allow"
assert got[("s3:PutObject", "arn:aws:s3:::retail-lake/raw/sales.csv")] == "Deny"
assert got[("s3:PutObject", "arn:aws:s3:::retail-lake/models/model.pkl")] == "Allow"
assert got[("s3:GetObject", "arn:aws:s3:::other-bucket/secret.csv")] == "Deny"
print(""" READ ROWS 2 AND 3 TOGETHER. The same action on the same
bucket is denied under raw/ and allowed under models/, because
a Deny statement scoped to one prefix beats an Allow scoped to
the bucket. Prefix-scoped policies are how a data lake keeps a
raw zone immutable while the rest stays writable""")
# Step 2: Add full S3 admin
print("\n now ADD a policy granting s3:* on everything:")
admin = {"name": "S3FullAccess",
"statements": [{"Effect": "Allow", "Action": ["s3:*"],
"Resource": ["*"]}]}
with_admin = f.POLICIES + [admin]
trace = {}
still = evaluate(with_admin, "s3:PutObject",
"arn:aws:s3:::retail-lake/raw/sales.csv", trace)
print(f" s3:PutObject on raw/ -> {still} ({trace['reason']})")
assert still == "Deny"
print(""" STILL DENIED, with full S3 admin attached. An explicit
Deny cannot be out-voted, out-numbered or out-scoped -- there
is no 'more specific allow wins' rule. To lift it you must
REMOVE the Deny.
This is the single most common IAM misunderstanding, and it is
also the feature: a Deny is how an organisation guarantees
something, rather than hoping nobody granted otherwise""")
# Step 3: Reverse the policy order
reversed_order = list(reversed(with_admin))
assert evaluate(reversed_order, "s3:PutObject",
"arn:aws:s3:::retail-lake/raw/sales.csv") == "Deny"
print("\n policy ORDER does not matter -- reversed, the answer is the same")
print(""" unlike a firewall rule list, IAM is not first-match.
Every statement is evaluated, then the three rules decide.
Say that and you have answered 'how does IAM resolve
conflicting policies?'""")
# Step 4: Compare a role with a user
print("\n a ROLE is not a user:")
print(f" {'':<16}{'user':<30}{'role'}")
for label, u, r in (
("credentials", "long-lived access key", "TEMPORARY, auto-rotated"),
("who assumes it", "a person", "a SERVICE or another principal"),
("in a notebook", "keys in a file <- BAD", "attached; no keys exist"),
("if leaked", "valid until revoked", "expires in minutes to hours")):
print(f" {label:<16}{u:<30}{r}")
print(""" a SageMaker notebook gets an EXECUTION ROLE, so no access
key is ever written to disk. That is why experiment 10 says
'attach IAM role' rather than 'paste your credentials', and
'I put my keys in the notebook' is the answer that loses the
marks""")
# Step 5: Compare least privilege with *:*
print("\n least privilege, as an exercise:")
over = {"name": "ItWorksNow",
"statements": [{"Effect": "Allow", "Action": ["*"],
"Resource": ["*"]}]}
tight = {"name": "TrainingJobOnly",
"statements": [
{"Effect": "Allow",
"Action": ["s3:GetObject"],
"Resource": ["arn:aws:s3:::retail-lake/train/*"]},
{"Effect": "Allow",
"Action": ["s3:PutObject"],
"Resource": ["arn:aws:s3:::retail-lake/models/*"]},
]}
probe = [("s3:GetObject", "arn:aws:s3:::retail-lake/train/x.csv"),
("s3:PutObject", "arn:aws:s3:::retail-lake/models/m.tar.gz"),
("iam:CreateUser", "*"),
("ec2:TerminateInstances", "*")]
print(f" {'action':<26}{'*:* policy':<14}{'scoped policy'}")
for a, r in probe:
print(f" {a:<26}{evaluate([over], a, r):<14}"
f"{evaluate([tight], a, r)}")
assert evaluate([over], "iam:CreateUser", "*") == "Allow"
assert evaluate([tight], "iam:CreateUser", "*") == "Deny"
print(""" both policies let the training job run. One of them also
lets it create IAM users and terminate every instance in the
account. '*:* made it work' is not a solution, it is a
postponed incident""")
# Step 6: Price the free tier
print("\n the free tier, and the three things that bill anyway:")
print(f" {'service':<22}{'free tier':<34}{'what still costs'}")
for svc, free, cost in (
("EC2", "750 hrs/month t2/t3.micro, 12 mo", "any larger instance"),
("S3", "5 GB Standard, 12 mo", "EGRESS to the internet"),
("RDS", "750 hrs/month db.t3.micro, 12 mo", "storage over 20 GB"),
("Lambda", "1M requests/month, ALWAYS free", "duration x memory"),
("SageMaker", "250 hrs notebook, 2 mo", "ENDPOINTS, billed hourly")):
print(f" {svc:<22}{free:<34}{cost}")
endpoint_month = f.EC2["m5.large"] * f.HOURS_PER_MONTH
print(f"\n a forgotten ml.m5.large endpoint costs about "
f"${endpoint_month:,.0f}/month")
assert 60 < endpoint_month < 90
print(""" THE ENDPOINT IS THE TRAP. A training job ends and stops
billing; an endpoint runs until you delete it, at hourly
rates, whether or not anything calls it. Every 'I got a
surprise AWS bill' story is a resource nobody switched off --
set a BUDGET ALARM on day one, before anything else""")
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 03_account_setup.md:
NOT RUN HERE
03_account_setup.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:
OUTPUT
Experiments 3 and 10 -- accounts, roles and IAM evaluation
the three rules, in order:
1. an EXPLICIT DENY anywhere wins -- always, unconditionally
2. otherwise an ALLOW that matches grants access
3. otherwise DENY -- the IMPLICIT DENY
the attached policies, evaluated:
action resource result why
s3:GetObject arn:aws:s3:::retail-lake/raw/sales.csv Allow Allow in DataScientistRead
s3:PutObject arn:aws:s3:::retail-lake/raw/sales.csv Deny EXPLICIT DENY in ProtectRawZone
s3:PutObject arn:aws:s3:::retail-lake/models/model.pkl Allow Allow in SageMakerExecution
s3:DeleteObject arn:aws:s3:::retail-lake/raw/sales.csv Deny EXPLICIT DENY in ProtectRawZone
s3:GetObject arn:aws:s3:::other-bucket/secret.csv Deny IMPLICIT DENY -- no statement matched
sagemaker:CreateTrainingJob arn:aws:sagemaker:*:*:training-job/x Allow Allow in SageMakerExecution
READ ROWS 2 AND 3 TOGETHER. The same action on the same
bucket is denied under raw/ and allowed under models/, because
a Deny statement scoped to one prefix beats an Allow scoped to
the bucket. Prefix-scoped policies are how a data lake keeps a
raw zone immutable while the rest stays writable
now ADD a policy granting s3:* on everything:
s3:PutObject on raw/ -> Deny (EXPLICIT DENY in ProtectRawZone)
STILL DENIED, with full S3 admin attached. An explicit
Deny cannot be out-voted, out-numbered or out-scoped -- there
is no 'more specific allow wins' rule. To lift it you must
REMOVE the Deny.
This is the single most common IAM misunderstanding, and it is
also the feature: a Deny is how an organisation guarantees
something, rather than hoping nobody granted otherwise
policy ORDER does not matter -- reversed, the answer is the same
unlike a firewall rule list, IAM is not first-match.
Every statement is evaluated, then the three rules decide.
Say that and you have answered 'how does IAM resolve
conflicting policies?'
a ROLE is not a user:
user role
credentials long-lived access key TEMPORARY, auto-rotated
who assumes it a person a SERVICE or another principal
in a notebook keys in a file <- BAD attached; no keys exist
if leaked valid until revoked expires in minutes to hours
a SageMaker notebook gets an EXECUTION ROLE, so no access
key is ever written to disk. That is why experiment 10 says
'attach IAM role' rather than 'paste your credentials', and
'I put my keys in the notebook' is the answer that loses the
marks
least privilege, as an exercise:
action *:* policy scoped policy
s3:GetObject Allow Allow
s3:PutObject Allow Allow
iam:CreateUser Allow Deny
ec2:TerminateInstances Allow Deny
both policies let the training job run. One of them also
lets it create IAM users and terminate every instance in the
account. '*:* made it work' is not a solution, it is a
postponed incident
the free tier, and the three things that bill anyway:
service free tier what still costs
EC2 750 hrs/month t2/t3.micro, 12 mo any larger instance
S3 5 GB Standard, 12 mo EGRESS to the internet
RDS 750 hrs/month db.t3.micro, 12 mo storage over 20 GB
Lambda 1M requests/month, ALWAYS free duration x memory
SageMaker 250 hrs notebook, 2 mo ENDPOINTS, billed hourly
a forgotten ml.m5.large endpoint costs about $70/month
THE ENDPOINT IS THE TRAP. A training job ends and stops
billing; an endpoint runs until you delete it, at hourly
rates, whether or not anything calls it. Every 'I got a
surprise AWS bill' story is a resource nobody switched off --
set a BUDGET ALARM on day one, before anything else
NOW ADD FULL S3 ADMIN
add a policy granting s3:* on *
s3:PutObject on raw/ -> Deny (EXPLICIT DENY in ProtectRawZone)
Still denied. An explicit Deny cannot be out-voted, out-numbered or out-scoped — there is no "more specific allow wins" rule. To lift it you must remove the Deny.
This is the single most common IAM misunderstanding, and it is also the feature: a Deny is how an organisation guarantees something rather than hoping nobody granted otherwise.
And policy order does not matter — reversed, the answer is identical. Unlike a firewall rule list, IAM is not first-match: every statement is evaluated, then the three rules decide.
Least privilege, made concrete:
| Action | *:* policy |
scoped policy |
|---|---|---|
s3:GetObject on train/ |
Allow | Allow |
s3:PutObject on models/ |
Allow | Allow |
iam:CreateUser |
Allow | Deny |
ec2:TerminateInstances |
Allow | Deny |
Both policies let the training job run. One of them also lets it create
IAM users and terminate every instance in the account. "*:* made it work"
is not a solution, it is a postponed incident.
RESULT
An explicit Deny on raw/ beats every Allow, even full S3 admin; policy order changes nothing; a scoped policy runs the job and refuses everything else.
Create and manage storage buckets, and upload and access datasets.
Store and list objects, and see that an object store has no directories, no rename and a bill for every version.
The procedure, on the console and CLI, 04_buckets.md:
The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:
THERE ARE NO DIRECTORIES
LIST prefix 'raw/' with delimiter '/':
objects at this level : (none)
common prefixes : ['raw/2026/']
raw/2026/01/sales.csv is ONE KEY containing three slashes. The console's
folder tree is drawn from common prefixes computed at list time. Delete
every object under a "folder" and the folder is gone, because it never
existed.
And a prefix scan is the only query an object store supports.
The procedure, on the console and CLI, 04_buckets.md:
# Experiment 4 -- create and manage storage buckets; upload and access datasets
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`04_storage.py`, which implements the key semantics and runs the pricing arithmetic**.
---
<!-- Step 1: Create a bucket and copy data, with the CLI -->
## The CLI, which is faster than the console
```bash
aws s3 mb s3://retail-lake-<your-unique-suffix> # names are GLOBAL
aws s3 cp sales.csv s3://retail-lake/raw/2026/01/
aws s3 cp ./data s3://retail-lake/raw/ --recursive
aws s3 ls s3://retail-lake/raw/ # by PREFIX
aws s3 ls s3://retail-lake/raw/ --recursive --human-readable --summarize
aws s3 sync ./local s3://retail-lake/curated/ # only what changed
aws s3 mv s3://retail-lake/a.csv s3://retail-lake/b.csv # COPY + DELETE
aws s3 rm s3://retail-lake/raw/old.csv
aws s3 presign s3://retail-lake/raw/sales.csv --expires-in 3600
```
**`s3 mv` is not a rename.** There is no rename. It copies the whole object
and deletes the original, and you are billed for both.
<!-- Step 2: Name it uniquely -->
## Bucket names are globally unique
Across every AWS customer on earth. `s3://data` was taken in 2006. Use
`<org>-<purpose>-<region>-<random>`.
<!-- Step 3: Turn on versioning and a lifecycle rule -->
## Versioning and lifecycle
```bash
aws s3api put-bucket-versioning --bucket retail-lake \
--versioning-configuration Status=Enabled
```
**Then set a lifecycle rule immediately**, or you pay for every version of
every object for ever:
```json
{"Rules": [{
"ID": "tier-and-expire",
"Filter": {"Prefix": "raw/"},
"Status": "Enabled",
"Transitions": [
{"Days": 30, "StorageClass": "STANDARD_IA"},
{"Days": 90, "StorageClass": "GLACIER"}
],
"NoncurrentVersionExpiration": {"NoncurrentDays": 30},
"AbortIncompleteMultipartUpload": {"DaysAfterInitiation": 7}
}]}
```
**That last rule matters more than it looks.** A failed multipart upload
leaves parts that are invisible in `s3 ls` and billed for ever. Many
mysterious S3 bills are abandoned multipart uploads.
<!-- Step 4: Block public access -->
## Blocking public access
```bash
aws s3api put-public-access-block --bucket retail-lake \
--public-access-block-configuration \
"BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"
```
**On by default since 2023**, and it should stay on. "Public S3 bucket" was
the single most common cause of data breaches for a decade. To share one
object, use a **presigned URL** — time-limited, no policy change.
<!-- Step 5: Encrypt it -->
## Encryption
Server-side encryption is on by default (SSE-S3). Use **SSE-KMS** when you
need an audit trail of who decrypted what, and accept that KMS charges per
request. `--sse aws:kms --sse-kms-key-id <arn>`.
The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:
"""Experiments 4, 5 and 6 -- object, block and file storage on the cloud.
The console click-paths for S3, EBS and EFS are in `04_buckets.md`,
`05_ebs.md` and `06_efs.md`, all marked as not run here -- there is no cloud
account.
What runs here is the SEMANTICS, which is what the exam is about: why an
object store has no directories, why you cannot rename, why a block volume
attaches to exactly one instance, and the pricing arithmetic that decides
which one you should have chosen.
"""
import fixtures as f
from objectstore import ObjectStore
def money(x):
return f"${x:,.2f}"
def main():
print(" Experiments 4, 5 and 6 -- object, block and file storage")
# Step 1: Experiment 4: list a bucket by prefix
# ================================================== experiment 4
print("\n --- experiment 4: buckets and objects")
s3 = ObjectStore("retail-lake")
keys = [
"raw/2026/01/sales.csv",
"raw/2026/02/sales.csv",
"raw/2026/03/sales.csv",
"curated/2026/sales.parquet",
"models/churn/model.tar.gz",
"README.md",
]
for k in keys:
s3.put(k, b"x" * 1024)
print(f" {len(s3.objects)} objects, {s3.total_bytes():,} bytes")
print("\n LIST with prefix 'raw/' and delimiter '/':")
plain, prefixes = s3.list("raw/", delimiter="/")
print(f" objects at this level : {plain or '(none)'}")
print(f" common prefixes : {prefixes}")
assert plain == [] and prefixes == ["raw/2026/"]
print(""" THERE IS NO DIRECTORY ANYWHERE. 'raw/2026/01/sales.csv' is
ONE KEY containing three slashes, and the console's folder
tree is drawn from COMMON PREFIXES computed at list time.
Delete every object under a 'folder' and the folder is gone,
because it never existed""")
print("\n LIST with prefix 'raw/' and NO delimiter:")
plain, prefixes = s3.list("raw/")
for k in plain:
print(f" {k}")
assert len(plain) == 3 and prefixes == []
print(""" without a delimiter you get every key under the prefix,
flat. That is the only query an object store supports: a
PREFIX SCAN. No WHERE, no index, no search by content""")
# Step 2: Rename an object
cost = s3.rename("README.md", "docs/README.md")
print(f"\n 'rename' README.md -> docs/README.md")
print(f" bytes read {cost['bytes_read']:,}, "
f"bytes written {cost['bytes_written']:,}, "
f"API calls {cost['requests']}")
assert "README.md" not in s3.objects and "docs/README.md" in s3.objects
print(""" THERE IS NO RENAME. It is a COPY plus a DELETE, so it
reads and writes the whole object. Renaming a 5 TB dataset
'to tidy up the folders' moves 10 TB and is billed for it --
and on a filesystem it would have been a metadata edit""")
# Step 3: Version it
print("\n versioning ON, then overwrite and delete:")
s3.versioning = True
s3.put("curated/2026/sales.parquet", b"y" * 2048)
s3.delete("curated/2026/sales.parquet")
hist = s3.versions["curated/2026/sales.parquet"]
current = s3.objects["curated/2026/sales.parquet"]
print(f" older versions kept : {len(hist)}")
print(f" current object : {current['class']}")
assert len(hist) == 2 and current["class"] == "DeleteMarker"
print(""" a DELETE with versioning on writes a DELETE MARKER; the
data is still there and still billed. That is the feature
(you can undelete) and the bill (you are paying for every
version of every object until a lifecycle rule removes
them)""")
# Step 4: Price the storage classes
print("\n storage classes -- 1 TB stored for a year, "
"retrieved once:")
gb = 1024
print(f" {'class':<22}{'storage/yr':>12}{'retrieve 1 TB':>15}"
f"{'total':>12}{'min days':>10}")
rows = {}
for cls in f.S3_STORAGE:
store = f.S3_STORAGE[cls] * gb * 12
retrieve = f.S3_RETRIEVAL[cls] * gb
rows[cls] = store + retrieve
print(f" {cls:<22}{money(store):>12}{money(retrieve):>15}"
f"{money(store + retrieve):>12}{f.S3_MIN_DAYS[cls]:>10}")
ratio = rows["Standard"] / rows["Glacier Deep Archive"]
store_only = f.S3_STORAGE["Standard"] / f.S3_STORAGE["Glacier Deep Archive"]
assert rows["Glacier Deep Archive"] < rows["Standard"] / 5
assert store_only > 2 * ratio
print(f""" READ THE TWO RATIOS SEPARATELY. Deep Archive STORAGE is
{store_only:.0f}x cheaper than Standard -- but add ONE retrieval a year
and the all-in saving falls to {ratio:.1f}x, because the retrieval
fee ({money(f.S3_RETRIEVAL['Glacier Deep Archive'] * gb)}) is larger than a year of its storage
({money(f.S3_STORAGE['Glacier Deep Archive'] * gb * 12)}). The headline discount is not the discount.
And it has a 180-DAY MINIMUM BILLING DURATION plus a
retrieval that takes up to 12 hours: delete an object after
10 days and you are billed for 180.
Storage class is a bet on your ACCESS PATTERN, and the
penalties are what make the bet real""")
# Step 5: Read it twice a month
print("\n the same 1 TB, retrieved TWICE A MONTH:")
print(f" {'class':<22}{'storage/yr':>12}{'retrieval/yr':>14}{'total':>12}")
freq = {}
for cls in ("Standard", "Standard-IA", "Glacier Instant"):
store = f.S3_STORAGE[cls] * gb * 12
retrieve = f.S3_RETRIEVAL[cls] * gb * 24
freq[cls] = store + retrieve
print(f" {cls:<22}{money(store):>12}{money(retrieve):>14}"
f"{money(store + retrieve):>12}")
assert freq["Standard"] < freq["Standard-IA"]
assert freq["Standard"] < freq["Glacier Instant"]
print(""" STANDARD IS NOW THE CHEAPEST. The 'cheap' tiers charge
per GB retrieved, and at two retrievals a month the retrieval
fee exceeds everything the storage discount saved.
Infrequent-access tiers are for data you genuinely do not
touch -- and 'we moved everything to IA to save money' is how
a bill goes UP""")
# Step 6: Price the egress
print("\n egress, which is the line item nobody predicts:")
print(f" {'transfer':<40}{'cost'}")
for label, amount in (("1 TB in from the internet", 0),
("1 TB out to the internet", gb * f.EGRESS_PER_GB),
("1 TB between AZs", gb * 0.01 * 2),
("1 TB S3 -> EC2, same region", 0)):
print(f" {label:<40}{money(amount)}")
month_store = f.S3_STORAGE["Standard"] * gb
one_egress = gb * f.EGRESS_PER_GB
months = one_egress / month_store
assert months > 3
print(f""" DOWNLOADING 1 TB ONCE ({money(one_egress)}) COSTS AS MUCH AS
STORING IT FOR {months:.1f} MONTHS ({money(month_store)}/month). Ingress is
free; egress is not. That asymmetry is the economic shape of
every cloud -- cheap to put data in, expensive to take it out
-- and it is the mechanism behind vendor lock-in: your data
is not held hostage, it is simply expensive to move.
It is also why 'move the compute to the data' survived from
Course 12 B into the cloud era: run the job in the region
holding the bucket and the transfer is free""")
# Step 7: Experiments 5 and 6: compare block, file and object storage
# ================================================== experiments 5, 6
print("\n --- experiments 5 and 6: block and file storage")
print(f"\n {'':<18}{'BLOCK (EBS)':<26}{'FILE (EFS)':<26}"
f"{'OBJECT (S3)'}")
for label, blk, fil, obj in (
("looks like", "a raw disk", "a mounted filesystem", "an HTTP API"),
("attached to", "ONE instance*", "MANY instances", "anything, anywhere"),
("access unit", "a 512 B block", "a file, byte ranges", "a whole object"),
("in-place edit", "yes", "yes", "NO -- rewrite it"),
("directories", "whatever the FS does", "real", "NONE -- prefixes"),
("latency", "sub-millisecond", "low millisecond", "tens of ms"),
("capacity", "provisioned, fixed", "elastic", "unlimited"),
("survives instance", "yes, if not root", "yes", "yes"),
("$/GB-month", f"{f.EBS_GP3_GB_MONTH:.3f}",
f"{f.EFS_STANDARD_GB_MONTH:.2f}",
f"{f.S3_STORAGE['Standard']:.3f}"),
):
print(f" {label:<18}{blk:<26}{fil:<26}{obj}")
print(" * EBS Multi-Attach exists for io1/io2 and needs a "
"cluster filesystem")
print("\n 1 TB for a month, by type:")
for label, price in (("EBS gp3", f.EBS_GP3_GB_MONTH),
("EFS Standard", f.EFS_STANDARD_GB_MONTH),
("S3 Standard", f.S3_STORAGE["Standard"])):
print(f" {label:<16}{money(price * gb):>10}")
ebs, efs, s3c = (p * gb for p in (f.EBS_GP3_GB_MONTH,
f.EFS_STANDARD_GB_MONTH,
f.S3_STORAGE["Standard"]))
assert efs > ebs > s3c
print(f""" EFS costs {efs / s3c:.0f}x what S3 does and {efs / ebs:.1f}x what EBS does,
and it is worth it precisely when several instances must
share a POSIX filesystem. Paying {efs / s3c:.0f}x for a dataset that one
batch job reads once is the mistake -- that dataset belongs
in S3.
Choose by ACCESS PATTERN, not by price: block for a database
or a boot disk, file for shared POSIX, object for everything
a data pipeline reads""")
# Step 8: Provision against consumption
print("\n provisioned against consumed, on a 1 TB EBS volume "
"holding 200 GB:")
print(f" EBS billed on PROVISIONED size : "
f"{money(f.EBS_GP3_GB_MONTH * gb)}/month")
print(f" S3 billed on STORED bytes : "
f"{money(f.S3_STORAGE['Standard'] * 200)}/month")
over = f.EBS_GP3_GB_MONTH * gb
used = f.S3_STORAGE["Standard"] * 200
assert over > used * 15
print(f""" a factor of {over / used:.0f}. EBS bills the volume you asked for,
empty or not; S3 bills the bytes you actually stored. An
over-provisioned volume is invisible waste, which is why
'just make it 1 TB to be safe' is an expensive habit""")
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 04_buckets.md:
NOT RUN HERE
04_buckets.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:
OUTPUT
Experiments 4, 5 and 6 -- object, block and file storage
--- experiment 4: buckets and objects
6 objects, 6,144 bytes
LIST with prefix 'raw/' and delimiter '/':
objects at this level : (none)
common prefixes : ['raw/2026/']
THERE IS NO DIRECTORY ANYWHERE. 'raw/2026/01/sales.csv' is
ONE KEY containing three slashes, and the console's folder
tree is drawn from COMMON PREFIXES computed at list time.
Delete every object under a 'folder' and the folder is gone,
because it never existed
LIST with prefix 'raw/' and NO delimiter:
raw/2026/01/sales.csv
raw/2026/02/sales.csv
raw/2026/03/sales.csv
without a delimiter you get every key under the prefix,
flat. That is the only query an object store supports: a
PREFIX SCAN. No WHERE, no index, no search by content
'rename' README.md -> docs/README.md
bytes read 1,024, bytes written 1,024, API calls 2
THERE IS NO RENAME. It is a COPY plus a DELETE, so it
reads and writes the whole object. Renaming a 5 TB dataset
'to tidy up the folders' moves 10 TB and is billed for it --
and on a filesystem it would have been a metadata edit
versioning ON, then overwrite and delete:
older versions kept : 2
current object : DeleteMarker
a DELETE with versioning on writes a DELETE MARKER; the
data is still there and still billed. That is the feature
(you can undelete) and the bill (you are paying for every
version of every object until a lifecycle rule removes
them)
storage classes -- 1 TB stored for a year, retrieved once:
class storage/yr retrieve 1 TB total min days
Standard $282.62 $0.00 $282.62 0
Standard-IA $153.60 $10.24 $163.84 30
One Zone-IA $122.88 $10.24 $133.12 30
Glacier Instant $49.15 $30.72 $79.87 90
Glacier Flexible $44.24 $10.24 $54.48 90
Glacier Deep Archive $12.17 $20.48 $32.65 180
READ THE TWO RATIOS SEPARATELY. Deep Archive STORAGE is
23x cheaper than Standard -- but add ONE retrieval a year
and the all-in saving falls to 8.7x, because the retrieval
fee ($20.48) is larger than a year of its storage
($12.17). The headline discount is not the discount.
And it has a 180-DAY MINIMUM BILLING DURATION plus a
retrieval that takes up to 12 hours: delete an object after
10 days and you are billed for 180.
Storage class is a bet on your ACCESS PATTERN, and the
penalties are what make the bet real
the same 1 TB, retrieved TWICE A MONTH:
class storage/yr retrieval/yr total
Standard $282.62 $0.00 $282.62
Standard-IA $153.60 $245.76 $399.36
Glacier Instant $49.15 $737.28 $786.43
STANDARD IS NOW THE CHEAPEST. The 'cheap' tiers charge
per GB retrieved, and at two retrievals a month the retrieval
fee exceeds everything the storage discount saved.
Infrequent-access tiers are for data you genuinely do not
touch -- and 'we moved everything to IA to save money' is how
a bill goes UP
egress, which is the line item nobody predicts:
transfer cost
1 TB in from the internet $0.00
1 TB out to the internet $92.16
1 TB between AZs $20.48
1 TB S3 -> EC2, same region $0.00
DOWNLOADING 1 TB ONCE ($92.16) COSTS AS MUCH AS
STORING IT FOR 3.9 MONTHS ($23.55/month). Ingress is
free; egress is not. That asymmetry is the economic shape of
every cloud -- cheap to put data in, expensive to take it out
-- and it is the mechanism behind vendor lock-in: your data
is not held hostage, it is simply expensive to move.
It is also why 'move the compute to the data' survived from
Course 12 B into the cloud era: run the job in the region
holding the bucket and the transfer is free
--- experiments 5 and 6: block and file storage
BLOCK (EBS) FILE (EFS) OBJECT (S3)
looks like a raw disk a mounted filesystem an HTTP API
attached to ONE instance* MANY instances anything, anywhere
access unit a 512 B block a file, byte ranges a whole object
in-place edit yes yes NO -- rewrite it
directories whatever the FS does real NONE -- prefixes
latency sub-millisecond low millisecond tens of ms
capacity provisioned, fixed elastic unlimited
survives instance yes, if not root yes yes
$/GB-month 0.080 0.30 0.023
* EBS Multi-Attach exists for io1/io2 and needs a cluster filesystem
1 TB for a month, by type:
EBS gp3 $81.92
EFS Standard $307.20
S3 Standard $23.55
EFS costs 13x what S3 does and 3.8x what EBS does,
and it is worth it precisely when several instances must
share a POSIX filesystem. Paying 13x for a dataset that one
batch job reads once is the mistake -- that dataset belongs
in S3.
Choose by ACCESS PATTERN, not by price: block for a database
or a boot disk, file for shared POSIX, object for everything
a data pipeline reads
provisioned against consumed, on a 1 TB EBS volume holding 200 GB:
EBS billed on PROVISIONED size : $81.92/month
S3 billed on STORED bytes : $4.60/month
a factor of 18. EBS bills the volume you asked for,
empty or not; S3 bills the bytes you actually stored. An
over-provisioned volume is invisible waste, which is why
'just make it 1 TB to be safe' is an expensive habit
THERE IS NO RENAME
'rename' README.md -> docs/README.md
bytes read 1,024, bytes written 1,024, API calls 2
Copy plus delete. Renaming a 5 TB dataset "to tidy the folders" moves 10 TB and is billed for it.
VERSIONING BILLS FOR EVERY VERSION
versioning ON, overwrite, then delete:
older versions kept : 2
current object : DeleteMarker
A delete writes a marker; the data is still there and still billed. Set the lifecycle rule when you enable versioning, not later.
Storage classes — 1 TB for a year, retrieved once:
| Class | Storage/yr | Retrieve | Total | Min days |
|---|---|---|---|---|
| Standard | $282.62 | $0.00 | $282.62 | 0 |
| Standard-IA | $153.60 | $10.24 | $163.84 | 30 |
| Glacier Instant | $49.15 | $30.72 | $79.87 | 90 |
| Deep Archive | $12.17 | $20.48 | $32.65 | 180 |
Read the two ratios separately. Deep Archive storage is 23× cheaper. Add one retrieval a year and the all-in saving falls to 8.7×, because the retrieval fee ($20.48) exceeds a whole year of its storage ($12.17). The headline discount is not the discount.
AND THE REVERSAL
The same 1 TB, retrieved twice a month:
| Class | Total/yr |
|---|---|
| Standard | $282.62 |
| Standard-IA | $399.36 |
| Glacier Instant | $786.43 |
Standard is now the cheapest. "We moved everything to IA to save money" is how a bill goes UP.
Egress:
| Transfer | Cost |
|---|---|
| 1 TB in | $0.00 |
| 1 TB out | $92.16 |
| 1 TB S3 → EC2, same region | $0.00 |
Downloading 1 TB once costs as much as storing it for 3.9 months. Ingress is free; egress is not — and that is the mechanism behind lock-in: your data is not held hostage, it is simply expensive to move.
RESULT
A listing by prefix finds no objects at raw/, only a common prefix; a rename is a copy and a delete; Deep Archive is 23× cheaper to store and 8.7× cheaper all-in.
Launch an instance and configure block storage (EBS).
Attach, format and mount a volume, and price block storage against file and object storage.
The procedure, on the console and CLI, 05_ebs.md:
BLOCK, FILE AND OBJECT
| EBS gp3 | EFS Standard | S3 Standard | |
|---|---|---|---|
| 1 TB/month | $81.92 | $307.20 | $23.55 |
And provisioned against consumed: a 1 TB EBS volume holding 200 GB bills $81.92 where S3 bills $4.60 — a factor of 18. "Just make it 1 TB to be safe" is an expensive habit.
The procedure, on the console and CLI, 05_ebs.md:
# Experiment 5 -- launch an instance and configure block storage (EBS)
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`04_storage.py`, which compares block, file and object semantics and cost**.
---
<!-- Step 1: Launch, attach, format and mount -->
## Launch, attach, format, mount
```bash
aws ec2 run-instances --image-id ami-xxxx --instance-type t3.micro \
--key-name mykey --security-group-ids sg-xxxx --subnet-id subnet-xxxx
aws ec2 create-volume --size 20 --volume-type gp3 \
--availability-zone us-east-1a
aws ec2 attach-volume --volume-id vol-xxxx --instance-id i-xxxx \
--device /dev/sdf
```
Then, on the instance:
```bash
lsblk # find it -- it appears as /dev/nvme1n1
sudo file -s /dev/nvme1n1 # "data" means NO FILESYSTEM yet
sudo mkfs -t xfs /dev/nvme1n1 # DESTROYS anything on it
sudo mkdir /data && sudo mount /dev/nvme1n1 /data
sudo blkid # get the UUID
echo 'UUID=<uuid> /data xfs defaults,nofail 0 2' | sudo tee -a /etc/fstab
```
<!-- Step 2: Avoid the three mistakes -->
## The three that go wrong
**1. `mkfs` on the wrong device.** `lsblk` first, every time. Formatting the
root volume ends the instance.
**2. Forgetting `/etc/fstab`.** The volume is not mounted after a reboot and
your application starts writing to the root disk instead — silently, until it
fills.
**3. `nofail` omitted.** If the volume is missing at boot, the instance hangs
in the boot sequence and you cannot SSH in to fix it. `nofail` turns a
catastrophe into a missing directory.
<!-- Step 3: Keep it in one zone -->
## An EBS volume is in ONE availability zone
You cannot attach `us-east-1a`'s volume to an instance in `us-east-1b`. To
move it: snapshot it (snapshots go to S3, which is regional), then create a
volume from the snapshot in the other AZ.
**That constraint is why "an EBS volume attaches to one instance" is really
"one instance, in one AZ"** — and it is why EFS exists.
<!-- Step 4: Choose a volume type -->
## Volume types, and how to choose
| Type | For | Notes |
|---|---|---|
| **gp3** | almost everything | IOPS and throughput are **independent of size** |
| gp2 | legacy | IOPS scale with size — the reason gp3 replaced it |
| io2 Block Express | a demanding database | expensive, and supports Multi-Attach |
| st1 | big sequential reads | HDD; terrible at random I/O |
| sc1 | cold archive on a disk | cheapest, slowest |
**gp3 is the default answer.** Under gp2 you would over-provision a volume
purely to buy IOPS; gp3 unbundled them.
<!-- Step 5: Snapshot it -->
## Snapshots
```bash
aws ec2 create-snapshot --volume-id vol-xxxx --description "before upgrade"
```
Snapshots are **incremental** — only changed blocks — and stored in S3.
Deleting an intermediate snapshot does not break later ones; AWS re-parents
the blocks. But snapshots of a **mounted, busy filesystem** may be
crash-consistent rather than clean: freeze or unmount for a database.
The procedure, on the console and CLI, 05_ebs.md:
NOT RUN HERE
05_ebs.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model for this experiment, 04_storage.py, is shown in full under Experiment 4, with what it printed.
RESULT
1 TB costs $81.92 a month on EBS; a 1 TB volume holding 200 GB bills $81.92 where S3 would bill $4.60.
Create and configure file storage on a cloud VM (EFS).
Create and mount a shared file system, and decide when it is worth its price.
The procedure, on the console and CLI, 06_efs.md:
THE PRICE
EFS costs 13× S3 and 3.8× EBS — worth it precisely when several instances must share a POSIX filesystem, and a mistake for a dataset one batch job reads once. The figures are the table under Experiment 5.
The procedure, on the console and CLI, 06_efs.md:
# Experiment 6 -- create and configure file storage on a cloud VM (EFS)
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`04_storage.py`, which prices EFS against EBS and S3 on the same terabyte**.
---
<!-- Step 1: Create and mount -->
## Create and mount
```bash
aws efs create-file-system --performance-mode generalPurpose \
--throughput-mode elastic --tags Key=Name,Value=shared-data
aws efs create-mount-target --file-system-id fs-xxxx \
--subnet-id subnet-xxxx --security-groups sg-xxxx
```
On each instance:
```bash
sudo apt install -y amazon-efs-utils
sudo mkdir /shared
sudo mount -t efs -o tls fs-xxxx:/ /shared
echo 'fs-xxxx:/ /shared efs _netdev,tls 0 0' | sudo tee -a /etc/fstab
```
**`_netdev` is not optional.** It tells systemd to wait for the network
before mounting. Without it the mount fails at boot, every time.
<!-- Step 2: Open the security group -->
## The security group rule everyone forgets
The EFS mount target needs an inbound rule allowing **NFS (TCP 2049)** from
the instances' security group. Without it the mount hangs — it does not fail
with a useful message, it hangs — and this is the single most common EFS
problem.
<!-- Step 3: Decide what EFS is for -->
## What EFS is actually for
Mount it on **many instances at once**, and they all see the same files with
POSIX semantics — locking, permissions, in-place writes. That is the whole
feature, and nothing else in the storage lineup offers it.
| Use it for | Do not use it for |
|---|---|
| a shared home directory across a fleet | a dataset one batch job reads once |
| shared model artefacts several workers read | a database's data files |
| a lift-and-shift app that expects a filesystem | anything a pipeline can read from S3 |
<!-- Step 4: Count the cost -->
## Cost, which is the reason to think twice
EFS Standard is about **$0.30/GB-month** — roughly 13x S3 Standard and
nearly 4x EBS gp3. Use **EFS Infrequent Access** with a lifecycle policy
(`--lifecycle-policy TransitionToIA=AFTER_30_DAYS`) and the cold portion
drops by about 90%.
**"We put the training data on EFS because it was easy to mount" is how a
storage bill triples.** That data belongs in S3.
<!-- Step 5: Know the equivalents -->
## Azure and GCP
| | AWS | Azure | GCP |
|---|---|---|---|
| Managed NFS | EFS | Azure Files (NFS/SMB) | Filestore |
| Protocol | NFSv4.1 | SMB **and** NFS | NFSv3 |
**Azure Files speaks SMB**, which matters if your workload is Windows —
that is the one real difference between the three.
The procedure, on the console and CLI, 06_efs.md:
NOT RUN HERE
06_efs.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model for this experiment, 04_storage.py, is shown in full under Experiment 4, with what it printed.
RESULT
EFS costs 13× S3 and 3.8× EBS: worth it when several instances share a POSIX file system.
Set up Jupyter Notebook or Colab on a cloud VM.
Run a notebook's cells in order and assert their outputs, and never publish it unauthenticated.
The procedure, on the console and CLI, 07_notebook.md:
WHAT RUNS
Cells executed and asserted in 01_vm_and_hosting.py:
In [1]: import pandas as pd; import fixtures as f
In [2]: df.shape -> (9, 19)
In [3]: df.groupby('region')['revenue'].sum() -> South 10360.0
In [4]: df['revenue'].sum() -> 12880.0
Four cells, executed in order, every output asserted. That is what a
notebook test looks like — papermill and nbconvert --execute do exactly
this in CI, and a notebook nobody executes in CI is a notebook that has
already drifted.
The procedure, on the console and CLI, 07_notebook.md:
# Experiment 7 -- set up Jupyter Notebook / Colab on a cloud VM
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`01_vm_and_hosting.py`, which executes notebook cells and asserts every output**.
---
<!-- Step 1: Start the notebook on the VM -->
## On the VM
```bash
sudo apt install -y python3-pip
pip install jupyterlab pandas scikit-learn matplotlib
jupyter lab --generate-config
jupyter lab password # set one, do not skip this
```
Then, **do not** open it to the internet. Use an SSH tunnel:
```bash
# on the VM
jupyter lab --no-browser --port=8888 --ip=127.0.0.1
# on your machine
ssh -N -L 8888:localhost:8888 ubuntu@<vm-ip>
# then browse to http://localhost:8888
```
<!-- Step 2: Avoid the mistake that matters -->
## ⚠ The mistake that matters
```bash
jupyter lab --ip=0.0.0.0 --allow-root --NotebookApp.token=''
```
**That publishes a root shell on the internet.** A Jupyter notebook executes
arbitrary code by design, so an unauthenticated notebook is not "an insecure
notebook" — it is a remote code execution endpoint. Scanners find these in
minutes; it is a standard way cloud accounts get used for cryptomining.
**Always: an SSH tunnel, or a managed notebook behind IAM.**
<!-- Step 3: Or use a managed notebook -->
## SageMaker Studio / Vertex Workbench instead
```bash
aws sagemaker create-notebook-instance \
--notebook-instance-name lab7 --instance-type ml.t3.medium \
--role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole
```
You get the tunnel, the authentication and the IAM role for free — **and no
access key is ever written to disk**, which is the real argument for it.
<!-- Step 4: Set the lifecycle configuration -->
## The lifecycle configuration that saves the money
```bash
#!/bin/bash
# attach as an OnStart lifecycle config
IDLE_TIME=3600
pip install -q jupyter-autoshutdown-extension || true
echo "auto-shutdown after ${IDLE_TIME}s idle" >> /var/log/lifecycle.log
```
**Colab disconnects after about 90 minutes idle and that is an annoyance. A
cloud notebook does not disconnect, and that is a bill.** An idle-shutdown
policy is the single most useful thing to configure on day one.
<!-- Step 5: Compare it with Colab -->
## Colab against a cloud notebook
| | Colab | Cloud notebook |
|---|---|---|
| Cost | free tier, then paid | the **instance**, hourly |
| Data | upload, or mount Drive | IAM role, no keys |
| Stops when | idle ~90 min | **never — you stop it** |
| State on stop | **lost** | kept on the volume |
| GPU | when available | the one you pay for |
| Private data | a policy question | inside your VPC |
**Colab is excellent for learning and wrong for anything confidential.** The
deciding question is whose infrastructure the data may sit on.
The procedure, on the console and CLI, 07_notebook.md:
NOT RUN HERE
07_notebook.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
THE MISTAKE THAT MATTERS
jupyter lab --ip=0.0.0.0 --allow-root --NotebookApp.token=''
That publishes a root shell on the internet. A notebook executes arbitrary code by design, so an unauthenticated one is not "an insecure notebook" — it is a remote code execution endpoint. Scanners find these in minutes.
Always an SSH tunnel, or a managed notebook behind IAM.
AND THE ROW THAT COSTS MONEY
| Colab | Cloud notebook | |
|---|---|---|
| Stops when | idle ~90 min | never — you stop it |
| State on stop | lost | kept on the volume |
An m5.xlarge notebook left running costs about $140/month. Colab disconnecting is an annoyance; a cloud notebook not disconnecting is a bill. Set an idle-shutdown lifecycle policy on day one.
The Python model for this experiment, 01_vm_and_hosting.py, is shown in full under Experiment 1, with what it printed.
RESULT
Four cells executed in order, every output asserted: South ₹10,360 and ₹12,880 in all.
Connect to cloud-hosted database services: RDS, BigQuery and Cosmos DB.
Query a managed database and a serverless warehouse, and see what a query costs.
The procedure, on the console and CLI, 08_cloud_db.md:
WHAT A QUERY COSTS
BigQuery, at $6.25 per TB scanned:
| Query | TB scanned | Cost |
|---|---|---|
SELECT * FROM events |
10.00 | $62.50 |
SELECT user_id FROM events |
0.40 | $2.50 |
SELECT user_id … WHERE dt = '…' |
0.02 | $0.12 |
The same question, 500× the price. Column projection and partition
pruning — Big Data Technologies' techniques, saving money here instead of time.
That is why SELECT * is a billing incident on a serverless warehouse and
merely rude on a server you already own.
The procedure, on the console and CLI, 08_cloud_db.md:
# Experiment 8 -- connect to cloud-hosted database services (RDS, BigQuery, Cosmos DB)
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`09_etl_warehouse.py`, which runs the same extract and the same warehouse queries locally**.
---
<!-- Step 1: Create a managed database -->
## RDS (managed PostgreSQL/MySQL)
```bash
aws rds create-db-instance \
--db-instance-identifier retail-db --db-instance-class db.t3.micro \
--engine postgres --allocated-storage 20 \
--master-username admin --manage-master-user-password \
--no-publicly-accessible --backup-retention-period 7
```
```bash
psql -h retail-db.xxxx.us-east-1.rds.amazonaws.com -U admin -d postgres
```
**`--no-publicly-accessible` is the right default.** Reach it from an EC2
instance in the same VPC, or through a bastion, or with SSM Session Manager.
A publicly reachable database with a weak password is found by scanners in
hours.
**And the security group needs an inbound rule for port 5432 from your
instance's security group** — group-to-group, not a CIDR. That is the fix for
almost every "connection timed out".
<!-- Step 2: Query BigQuery -->
## BigQuery
```bash
bq mk --dataset retail
bq load --autodetect --source_format=CSV retail.sales gs://bucket/sales.csv
bq query --use_legacy_sql=false --dry_run \
'SELECT region, SUM(revenue) FROM retail.sales GROUP BY region'
```
**`--dry_run` prints the bytes the query would scan without running it, and
therefore what it would cost.** Run it before every large query. It is the
single most valuable BigQuery habit.
```bash
bq query --use_legacy_sql=false --maximum_bytes_billed=1000000000 '...'
```
**`--maximum_bytes_billed` is a seatbelt**: the query fails rather than
billing more than you said. Set it in every automated job.
<!-- Step 3: Try Cosmos DB -->
## Cosmos DB
```bash
az cosmosdb create --name retail-cosmos --resource-group rg \
--kind GlobalDocumentDB --default-consistency-level Session
az cosmosdb sql database create --account-name retail-cosmos \
--resource-group rg --name retail
```
**The consistency level is the interesting choice**, and it is Course 10's
CAP discussion made into a dropdown:
| Level | Guarantee | Cost |
|---|---|---|
| Strong | linearizable | highest RU, single region write |
| **Bounded staleness** | lag bounded by time or versions | high |
| **Session** | your own writes are read back | **the default, and usually right** |
| Consistent prefix | order preserved, some lag | low |
| Eventual | none | **lowest RU** |
**Cosmos bills in Request Units**, and stronger consistency costs more RUs
per read. That is CAP with a price tag attached.
<!-- Step 4: Compare managed with self-hosted -->
## Managed against self-hosted
| | Self-hosted on EC2 | Managed (RDS) |
|---|---|---|
| Patching | you | automatic, in a window |
| Backups | you write them | continuous, point-in-time |
| Failover | you build it | Multi-AZ, ~60 s |
| Version upgrade | your weekend | a click, with a restart |
| Root/superuser | yes | **no** |
| Cost | instance only | instance + ~20-30% |
**"No superuser" is the row that surprises people.** Some extensions, some
`ALTER SYSTEM` settings and any filesystem access are unavailable. If your
application needs those, managed is not an option — and knowing that before
you migrate is the point.
The procedure, on the console and CLI, 08_cloud_db.md:
NOT RUN HERE
08_cloud_db.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model for this experiment, 09_etl_warehouse.py, is shown in full under Experiment 9, with what it printed.
RESULT
The same question costs $62.50 as SELECT * and $0.12 with a column list and a partition filter.
Build a batch ETL pipeline: extract from an operational database, transform, and load a warehouse.
Run the pipeline end to end, with an audit trail at every step, and check South against three other engines.
The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:
THE PIPELINE, WITH AN AUDIT TRAIL
| Step | Rows |
|---|---|
| extracted | 11 |
| after dedup | 10 |
| dropped, null region | 1 |
| loaded | 9 |
11 in, 9 out, and the pipeline can say where the other two went. A transformation that silently drops rows is worse than one that fails: the numbers still look plausible. Every ETL job should emit these counts, and a monitoring rule should alarm when the drop rate moves.
The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:
"""Experiments 8, 9 and 12 -- cloud databases, a batch ETL pipeline, and
loading a cloud data warehouse.
AWS RDS, BigQuery and Cosmos DB need an account, so `08_cloud_db.md` and
`12_etl_to_warehouse.md` carry the console and CLI steps, marked NOT
EXECUTED.
What runs is the pipeline itself: extract from a real relational source
(SQLite, standing in for RDS), transform, and load into a real columnar
warehouse (DuckDB, standing in for Redshift/BigQuery). The transformations,
the row counts and the money all reconcile -- and they reconcile against
Course 11 and Course 12 B, which used the same nine facts.
The pricing arithmetic at the end is the part of "cloud data warehouse" that
is genuinely different from Course 5's DBMS, and it is what gets examined.
"""
import os
import sqlite3
import tempfile
import duckdb
import fixtures as f
SRC_COLUMNS = ["order_id", "date_key", "store", "region", "product",
"category", "qty", "list_price", "unit_cost"]
def money(x):
return f"${x:,.2f}"
def extract(db_path):
"""E -- read from the operational database. Never transform here."""
con = sqlite3.connect(db_path)
con.execute("""CREATE TABLE orders (
order_id INTEGER PRIMARY KEY, date_key TEXT, store TEXT, region TEXT,
product TEXT, category TEXT, qty INTEGER,
list_price REAL, unit_cost REAL)""")
rows = []
for i, (_, r) in enumerate(f.SALES_DF.iterrows()):
rows.append((i + 1, r["date_key"], r["store"], r["region"],
r["product"], r["category"], int(r["qty"]),
float(r["list_price"]), float(r["unit_cost"])))
# deliberately dirty: one duplicate and one null region
rows.append((10, rows[0][1], rows[0][2], rows[0][3], rows[0][4],
rows[0][5], rows[0][6], rows[0][7], rows[0][8]))
rows.append((11, "D4", "Guntur", None, "Tea 500g", "Grocery",
3, 210.0, 150.0))
con.executemany(
"INSERT INTO orders VALUES (?,?,?,?,?,?,?,?,?)", rows)
con.commit()
out = con.execute(
f"SELECT {', '.join(SRC_COLUMNS)} FROM orders").fetchall()
con.close()
return out
def transform(rows, log):
"""T -- and every step records what it dropped, or it is not auditable."""
log["extracted"] = len(rows)
seen, deduped = set(), []
for r in rows:
fingerprint = r[1:] # everything but the surrogate id
if fingerprint in seen:
continue
seen.add(fingerprint)
deduped.append(r)
log["after_dedup"] = len(deduped)
clean = [r for r in deduped if r[3] is not None]
log["dropped_null_region"] = len(deduped) - len(clean)
enriched = []
for r in clean:
d = dict(zip(SRC_COLUMNS, r))
d["revenue"] = d["qty"] * d["list_price"]
d["cost"] = d["qty"] * d["unit_cost"]
d["profit"] = d["revenue"] - d["cost"]
d["quarter"] = "Q1" if d["date_key"] in ("D1", "D2") else "Q2"
enriched.append(d)
log["loaded"] = len(enriched)
return enriched
def main():
print(" Experiments 8, 9 and 12 -- cloud DB, batch ETL, warehouse load")
tmp = tempfile.mkdtemp(prefix="cloud09_")
db = os.path.join(tmp, "orders.db")
# Step 1: Extract from the operational database
raw = extract(db)
print(f"\n EXTRACT from the operational database (SQLite as RDS):")
print(f" {len(raw)} rows, including one duplicate and one null region")
assert len(raw) == 11
# Step 2: Transform, with an audit trail
log = {}
clean = transform(raw, log)
print(f"\n TRANSFORM, with an audit trail at every step:")
print(f" {'step':<26}{'rows':>6}")
for step in ("extracted", "after_dedup", "dropped_null_region", "loaded"):
print(f" {step:<26}{log[step]:>6}")
assert log == {"extracted": 11, "after_dedup": 10,
"dropped_null_region": 1, "loaded": 9}
print(""" 11 in, 9 out, and the pipeline can SAY WHERE THE OTHER TWO
WENT. A transformation that silently drops rows is worse than
one that fails: the numbers still look plausible.
Every ETL job should emit these counts, and a monitoring rule
should alarm when the drop rate moves""")
# Step 3: Load into the warehouse
con = duckdb.connect()
con.execute("""CREATE TABLE fact_sales (
order_id INTEGER, date_key VARCHAR, store VARCHAR, region VARCHAR,
product VARCHAR, category VARCHAR, qty INTEGER,
list_price DOUBLE, unit_cost DOUBLE,
revenue DOUBLE, cost DOUBLE, profit DOUBLE, quarter VARCHAR)""")
con.executemany(
"INSERT INTO fact_sales VALUES (" + ",".join(["?"] * 13) + ")",
[[r[c] for c in ("order_id", "date_key", "store", "region", "product",
"category", "qty", "list_price", "unit_cost",
"revenue", "cost", "profit", "quarter")]
for r in clean])
n, rev = con.execute(
"SELECT COUNT(*), SUM(revenue) FROM fact_sales").fetchone()
print(f"\n LOAD into the warehouse (DuckDB as Redshift/BigQuery):")
print(f" {n} rows, revenue {money(rev)}")
assert n == 9 and rev == f.total_revenue()
print(f""" {money(rev)} is Course 11's total, Course 12 B's Hive total and
Course 12 B's Spark total. FOUR engines now agree on the same
nine facts, and the suite fails if any of them drifts""")
rows = con.execute("""
SELECT region, SUM(revenue) rev, SUM(profit) prof
FROM fact_sales GROUP BY region ORDER BY rev DESC""").fetchall()
print(f"\n {'region':<10}{'revenue':>12}{'profit':>10}")
for region, r, p in rows:
print(f" {region:<10}{money(r):>12}{money(p):>10}")
assert dict((r[0], r[1]) for r in rows)["South"] == 10360.0
# Step 4: Set ETL against ELT
print("\n ETL against ELT, which is the modern distinction:")
print(f" {'':<16}{'ETL':<34}{'ELT'}")
for label, etl, elt in (
("transform runs", "on a separate compute box", "IN the warehouse"),
("lands in the DW", "clean data only", "RAW data, then transformed"),
("re-run a change", "re-extract from source", "re-run SQL on raw"),
("needs", "an ETL server or Glue", "a warehouse that scales"),
("source load", "one read", "one read"),
("suits", "limited warehouse capacity", "cheap elastic compute")):
print(f" {label:<16}{etl:<34}{elt}")
print(""" ELT WON BECAUSE WAREHOUSE COMPUTE GOT CHEAP AND ELASTIC.
Landing raw data means a transformation bug is fixed by
re-running SQL rather than re-extracting from a production
database that may no longer hold the old rows -- which is
exactly the DELETE problem Course 12 B found in Sqoop""")
# Step 5: See what makes a cloud warehouse different
print("\n what makes a cloud DW different from Course 5's RDBMS:")
print(f" {'':<22}{'RDBMS (Course 5)':<28}{'cloud DW'}")
for label, rdbms, dw in (
("storage layout", "ROW", "COLUMNAR"),
("workload", "many small transactions", "few huge scans"),
("indexes", "central to performance", "usually none"),
("scaling", "a bigger box", "add nodes / serverless"),
("compute & storage", "coupled", "SEPARATED"),
("billed on", "the box, hourly", "BYTES SCANNED or node-hours"),
("a bad query costs", "time", "TIME AND MONEY")):
print(f" {label:<22}{rdbms:<28}{dw}")
# Step 6: Price the queries
print("\n BigQuery on-demand, at "
f"${f.BIGQUERY_PER_TB_SCANNED:.2f} per TB SCANNED:")
print(f" {'query':<44}{'TB scanned':>12}{'cost':>10}")
scenarios = [
("SELECT * FROM events ", 10.0),
("SELECT user_id FROM events ", 0.4),
("SELECT user_id ... WHERE dt = '2026-08-01'", 0.02),
]
costs = {}
for label, tb in scenarios:
c = tb * f.BIGQUERY_PER_TB_SCANNED
costs[label.strip()] = c
print(f" {label:<44}{tb:>12.2f}{money(c):>10}")
full = costs["SELECT * FROM events"]
pruned = costs["SELECT user_id ... WHERE dt = '2026-08-01'"]
assert full / pruned == 500
print(f""" THE SAME QUESTION, {full / pruned:.0f}x THE PRICE. Selecting one
column instead of all reads a fraction of the bytes
(Course 12 B's column projection), and a partition filter
removes almost all of the rest (Course 12 B's partition
pruning).
In Course 12 B those techniques saved TIME. Here they save
MONEY, on the same mechanism -- which is why 'SELECT *' is a
billing incident on a serverless warehouse and merely rude on
a server you already own""")
print(f"\n Redshift, at ${f.REDSHIFT_RA3_XLPLUS_HOUR:.3f} per node-hour:")
for nodes in (2, 4, 8):
m = nodes * f.REDSHIFT_RA3_XLPLUS_HOUR * f.HOURS_PER_MONTH
print(f" {nodes} nodes: {money(m):>12}/month, "
f"{'queries are free at the margin'}")
two = 2 * f.REDSHIFT_RA3_XLPLUS_HOUR * f.HOURS_PER_MONTH
breakeven_tb = two / f.BIGQUERY_PER_TB_SCANNED
print(f"""
break-even: {breakeven_tb:,.0f} TB scanned per month
BELOW that, on-demand BigQuery is cheaper and you pay nothing
when idle. ABOVE it, a provisioned cluster is cheaper and an
extra query costs nothing at the margin.
That is the whole provisioned-against-serverless decision,
and it is a calculation rather than a preference""")
assert 250 < breakeven_tb < 260
con.close()
os.remove(db)
os.rmdir(tmp)
if __name__ == "__main__":
main()
The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:
OUTPUT
Experiments 8, 9 and 12 -- cloud DB, batch ETL, warehouse load
EXTRACT from the operational database (SQLite as RDS):
11 rows, including one duplicate and one null region
TRANSFORM, with an audit trail at every step:
step rows
extracted 11
after_dedup 10
dropped_null_region 1
loaded 9
11 in, 9 out, and the pipeline can SAY WHERE THE OTHER TWO
WENT. A transformation that silently drops rows is worse than
one that fails: the numbers still look plausible.
Every ETL job should emit these counts, and a monitoring rule
should alarm when the drop rate moves
LOAD into the warehouse (DuckDB as Redshift/BigQuery):
9 rows, revenue $12,880.00
$12,880.00 is Course 11's total, Course 12 B's Hive total and
Course 12 B's Spark total. FOUR engines now agree on the same
nine facts, and the suite fails if any of them drifts
region revenue profit
South $10,360.00 $2,760.00
North $2,520.00 $765.00
ETL against ELT, which is the modern distinction:
ETL ELT
transform runs on a separate compute box IN the warehouse
lands in the DW clean data only RAW data, then transformed
re-run a change re-extract from source re-run SQL on raw
needs an ETL server or Glue a warehouse that scales
source load one read one read
suits limited warehouse capacity cheap elastic compute
ELT WON BECAUSE WAREHOUSE COMPUTE GOT CHEAP AND ELASTIC.
Landing raw data means a transformation bug is fixed by
re-running SQL rather than re-extracting from a production
database that may no longer hold the old rows -- which is
exactly the DELETE problem Course 12 B found in Sqoop
what makes a cloud DW different from Course 5's RDBMS:
RDBMS (Course 5) cloud DW
storage layout ROW COLUMNAR
workload many small transactions few huge scans
indexes central to performance usually none
scaling a bigger box add nodes / serverless
compute & storage coupled SEPARATED
billed on the box, hourly BYTES SCANNED or node-hours
a bad query costs time TIME AND MONEY
BigQuery on-demand, at $6.25 per TB SCANNED:
query TB scanned cost
SELECT * FROM events 10.00 $62.50
SELECT user_id FROM events 0.40 $2.50
SELECT user_id ... WHERE dt = '2026-08-01' 0.02 $0.12
THE SAME QUESTION, 500x THE PRICE. Selecting one
column instead of all reads a fraction of the bytes
(Course 12 B's column projection), and a partition filter
removes almost all of the rest (Course 12 B's partition
pruning).
In Course 12 B those techniques saved TIME. Here they save
MONEY, on the same mechanism -- which is why 'SELECT *' is a
billing incident on a serverless warehouse and merely rude on
a server you already own
Redshift, at $1.086 per node-hour:
2 nodes: $1,585.56/month, queries are free at the margin
4 nodes: $3,171.12/month, queries are free at the margin
8 nodes: $6,342.24/month, queries are free at the margin
break-even: 254 TB scanned per month
BELOW that, on-demand BigQuery is cheaper and you pay nothing
when idle. ABOVE it, a provisioned cluster is cheaper and an
extra query costs nothing at the margin.
That is the whole provisioned-against-serverless decision,
and it is a calculation rather than a preference
This experiment has no console procedure: SQLite stands in for the operational database (RDS) and DuckDB for the warehouse, and the pipeline between them runs end to end. Experiment 12's procedure does the same with Glue and Redshift or BigQuery.
THE FOUR-ENGINE CHECK
| Region | Revenue | Profit | Margin |
|---|---|---|---|
| South | 10,360 | 2,760 | 26.64% |
| North | 2,520 | 765 | 30.36% |
₹12,880 total, ₹10,360 for South — Business Intelligence Tools' DAX, Big Data Technologies' Hive, Big Data Technologies' Spark and this. Four engines, one set of nine facts.
And the break-even: Redshift at $1.086/node-hour: 2 nodes is $1,585.56/month. Break-even against on-demand BigQuery: about 254 TB scanned per month. Below that, serverless is cheaper and costs nothing when idle. Above it, a cluster is cheaper and an extra query is free at the margin. A calculation, not a preference.
RESULT
11 rows in, 9 loaded, and the pipeline says where the other two went; South is ₹10,360, as in DAX, Hive and Spark.
Launch a SageMaker notebook, and attach an IAM role and an S3 bucket.
Give the notebook a role rather than keys, and know what an idle endpoint costs.
The procedure, on the console and CLI, 10_sagemaker_notebook.md:
A ROLE IS NOT A USER
| User | Role | |
|---|---|---|
| Credentials | long-lived access key | temporary, auto-rotated |
| In a notebook | keys in a file — bad | attached; no keys exist |
| If leaked | valid until revoked | expires in minutes to hours |
A SageMaker notebook gets an execution role, so no access key is ever written to disk. That is why experiment 10 says "attach IAM role" rather than "paste your credentials", and "I put my keys in the notebook" is the answer that loses the marks.
The procedure, on the console and CLI, 10_sagemaker_notebook.md:
# Experiment 10 -- launch a SageMaker notebook, attach an IAM role and an S3 bucket
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`03_iam_and_account.py`, which implements IAM's evaluation algorithm and proves what the role can do**.
---
<!-- Step 1: Create the role first -->
## The role first, then the notebook
```bash
aws iam create-role --role-name SageMakerExecutionRole \
--assume-role-policy-document '{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {"Service": "sagemaker.amazonaws.com"},
"Action": "sts:AssumeRole"}]}'
```
**That document is the TRUST POLICY, and it is not the permissions.** It says
*who may become this role*; a separate permissions policy says *what the role
may then do*. Confusing the two is the commonest IAM error after the explicit
Deny.
```bash
aws iam put-role-policy --role-name SageMakerExecutionRole \
--policy-name S3Scoped --policy-document '{
"Version": "2012-10-17",
"Statement": [
{"Effect": "Allow", "Action": ["s3:GetObject"],
"Resource": "arn:aws:s3:::retail-lake/train/*"},
{"Effect": "Allow", "Action": ["s3:PutObject"],
"Resource": "arn:aws:s3:::retail-lake/models/*"}]}'
```
**Two statements, two prefixes, two actions.** Not `s3:*` on `*`. The
runnable half shows both policies letting the training job succeed, and only
one of them also permitting `iam:CreateUser`.
<!-- Step 2: Then the notebook -->
## Then the notebook
```bash
aws sagemaker create-notebook-instance \
--notebook-instance-name lab10 \
--instance-type ml.t3.medium \
--role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole \
--volume-size-in-gb 20
aws sagemaker describe-notebook-instance --notebook-instance-name lab10
aws sagemaker create-presigned-notebook-instance-url \
--notebook-instance-name lab10
```
Inside the notebook, **no credentials exist anywhere**:
```python
import boto3, sagemaker
session = sagemaker.Session()
role = sagemaker.get_execution_role() # reads the ATTACHED role
bucket = "retail-lake"
session.upload_data("train.csv", bucket=bucket, key_prefix="train")
```
<!-- Step 3: Avoid what not to do -->
## ⚠ What not to do
```python
boto3.client("s3", aws_access_key_id="AKIA...", aws_secret_access_key="...")
```
**Keys in a notebook end up in git.** GitHub scans for them and AWS
quarantines the account, which is the good outcome; the bad outcome is a
cryptomining bill. The execution role exists precisely so this is never
necessary.
<!-- Step 4: Stop it when you are done -->
## Stop it when you are done
```bash
aws sagemaker stop-notebook-instance --notebook-instance-name lab10
aws sagemaker delete-notebook-instance --notebook-instance-name lab10
```
**Stopping keeps the EBS volume (and its charge); deleting removes
everything.** A stopped notebook costs only storage; a running one costs the
instance whether or not the browser tab is open.
The procedure, on the console and CLI, 10_sagemaker_notebook.md:
NOT RUN HERE
10_sagemaker_notebook.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
THE ENDPOINT TRAP
A forgotten ml.m5.large endpoint costs about $70/month. A training job
ends and stops billing; an endpoint runs until you delete it, at hourly
rates, whether or not anything calls it.
Set a budget alarm on day one, before anything else.
The Python model for this experiment, 03_iam_and_account.py, is shown in full under Experiment 3, with what it printed.
RESULT
A role's credentials are temporary and never written to disk; a forgotten endpoint costs about $70 a month.
Build a classification or regression model on a managed ML platform.
Train a model, quote the dummy first, save the artefact, and price the instance.
The procedure, on the console and CLI, 11_sagemaker_train.md:
The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:
QUOTE THE DUMMY FIRST, ALWAYS
| Model | Accuracy | F1 | AUC |
|---|---|---|---|
DummyClassifier |
0.8433 | 0.0000 | 0.5000 |
| GradientBoosting | 0.9467 | 0.8095 | 0.9029 |
94.67% sounds excellent until you see 84.33% for predicting "never churns". The real gain is 10.3 percentage points, and the F1 of 0.8095 against 0.0000 is what shows the model found anything.
Machine Learning's argument, and it does not stop being true because the model trained on somebody else's computer.
The procedure, on the console and CLI, 11_sagemaker_train.md:
# Experiment 11 -- build a classification/regression model on a managed ML platform
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`11_train_and_automl.py`, which trains the same model and prices the instance choice**.
---
<!-- Step 1: Call the SDK -->
## The SDK call
```python
from sagemaker.sklearn.estimator import SKLearn
estimator = SKLearn(
entry_point="train.py",
role=role,
instance_type="ml.m5.xlarge",
instance_count=1,
framework_version="1.2-1",
hyperparameters={"n_estimators": 100, "max_depth": 5},
output_path=f"s3://{bucket}/models/",
)
estimator.fit({"train": f"s3://{bucket}/train/",
"validation": f"s3://{bucket}/validation/"})
```
<!-- Step 2: See what makes it managed -->
## The three things that make it a MANAGED job
1. **`entry_point="train.py"`** — your ordinary scikit-learn script, run
inside a container SageMaker builds. The algorithm is unchanged.
2. **`instance_type`** — billed per second, spun up for the job and destroyed
after. This is the actual product.
3. **`output_path`** — the artefact lands in S3. Training and serving are
separate systems joined by one file.
<!-- Step 3: Follow the train.py contract -->
## The `train.py` contract
```python
import argparse, os, joblib, pandas as pd
from sklearn.ensemble import GradientBoostingClassifier
parser = argparse.ArgumentParser()
parser.add_argument("--n_estimators", type=int, default=100)
parser.add_argument("--train", default=os.environ["SM_CHANNEL_TRAIN"])
parser.add_argument("--model-dir", default=os.environ["SM_MODEL_DIR"])
args = parser.parse_args()
df = pd.read_csv(os.path.join(args.train, "train.csv"))
X, y = df.drop(columns=["target"]), df["target"]
model = GradientBoostingClassifier(n_estimators=args.n_estimators).fit(X, y)
joblib.dump(model, os.path.join(args.model_dir, "model.joblib"))
```
**Hyperparameters arrive as command-line arguments; channels arrive as
environment variables; the model must be written to `SM_MODEL_DIR`.** Those
three conventions are the entire interface, and getting `SM_MODEL_DIR` wrong
is why a job "succeeds" and produces no artefact.
<!-- Step 4: Train on spot -->
## Spot training
```python
estimator = SKLearn(..., use_spot_instances=True,
max_run=3600, max_wait=7200,
checkpoint_s3_uri=f"s3://{bucket}/checkpoints/")
```
**Up to 70% cheaper, and interruptible.** `max_wait` must exceed `max_run` to
leave room for interruptions, and without `checkpoint_s3_uri` an interrupted
job restarts from zero. For a 20-minute job spot is free money; for a 20-hour
job without checkpoints it is a trap.
<!-- Step 5: Choose the instance -->
## Instance choice, which is answered by the algorithm
| Algorithm | Instance | Why |
|---|---|---|
| scikit-learn, XGBoost on tabular | **m5 / c5** | no GPU code path exists |
| Deep learning, training | p3 / p4d / g5 | dense matrix multiplication |
| Deep learning, inference | g4dn / inf1 | cheaper per prediction |
| Anything, if the data fits in RAM | the smallest that fits | you are paying for RAM |
**Gradient boosting on tabular data does not use a GPU.** A `p4d.24xlarge`
would run this model at the same speed for roughly 170 times the price — a
figure the runnable half computes.
The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:
"""Experiments 11 and 14 -- build a model on a managed ML platform, and use
an AutoML service.
SageMaker, Azure ML Studio and Vertex AI need an account, so
`11_sagemaker_train.md` and `14_automl.md` carry the console steps and the
SDK calls, marked as not run here.
What runs here is the model, on the same scikit-learn Course 12 A used --
because the ALGORITHM is not what the cloud changes. What the cloud changes
is the packaging: where the data comes from, where the artefact goes, and
what it costs. Those three are modelled explicitly.
The AutoML half is a genuine (small) AutoML: a real search over real models
with real cross-validation, so the leaderboard is measured. That makes the
point AutoML marketing does not -- what it actually does, and what it costs.
"""
import io
import json
import os
import tempfile
import time
import joblib
import numpy as np
from sklearn.datasets import make_classification
from sklearn.dummy import DummyClassifier
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score, roc_auc_score
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier
import fixtures as f
SEED = 42
MODEL_PATH = os.path.join(tempfile.gettempdir(), "cloud13b_model.joblib")
def churn_data(n=1200):
"""A churn-shaped dataset with a KNOWN base rate, as in Course 12 A."""
X, y = make_classification(
n_samples=n, n_features=10, n_informative=5, n_redundant=2,
weights=[0.85, 0.15], flip_y=0.02, class_sep=1.1,
random_state=SEED)
return X, y
def train_model(X_train, y_train):
"""The 'training job'. On SageMaker this is a container; here it is a call."""
pipe = Pipeline([("scale", StandardScaler()),
("clf", GradientBoostingClassifier(random_state=SEED))])
pipe.fit(X_train, y_train)
return pipe
def main():
print(" Experiments 11 and 14 -- training and AutoML on a managed platform")
# Step 1: Make the data
X, y = churn_data()
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=SEED)
base_rate = y.mean()
print(f"\n dataset: {len(X):,} rows, {X.shape[1]} features, "
f"base rate {base_rate:.2%} positive")
assert 0.13 < base_rate < 0.17
# Step 2: Experiment 11: run the training job
print("\n --- experiment 11: the training job")
t0 = time.perf_counter()
model = train_model(X_train, y_train)
train_seconds = time.perf_counter() - t0
pred = model.predict(X_test)
proba = model.predict_proba(X_test)[:, 1]
acc = accuracy_score(y_test, pred)
f1 = f1_score(y_test, pred)
auc = roc_auc_score(y_test, proba)
dummy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
dummy_acc = accuracy_score(y_test, dummy.predict(X_test))
print(f" {'model':<26}{'accuracy':>10}{'F1':>8}{'AUC':>8}")
print(f" {'DummyClassifier':<26}{dummy_acc:>10.4f}"
f"{0.0:>8.4f}{0.5:>8.4f}")
print(f" {'GradientBoosting':<26}{acc:>10.4f}{f1:>8.4f}{auc:>8.4f}")
assert acc > dummy_acc and auc > 0.85
print(f""" quote the DUMMY FIRST, always. {acc:.2%} sounds excellent
until you see that predicting 'never churns' scores {dummy_acc:.2%}
-- the real gain is {acc - dummy_acc:.1%} points of accuracy, and it is the
F1 of {f1:.4f} against 0.0000 that shows the model found anything
at all. This is Course 12 A's argument, and it does not stop
being true because the model trained on somebody else's
computer""")
# Step 3: Save the artefact, and reload it
joblib.dump(model, MODEL_PATH)
size = os.path.getsize(MODEL_PATH)
reloaded = joblib.load(MODEL_PATH)
assert (reloaded.predict(X_test) == pred).all()
print(f"\n model artefact: {size:,} bytes, reloads and predicts "
f"identically")
print(""" THAT FILE IS THE DELIVERABLE. A SageMaker training job
writes exactly this to s3://bucket/models/, and the deploy
step reads it back. Training and serving are separate
systems joined by one artefact in object storage -- which is
why the IAM role in experiment 10 needs s3:PutObject on
models/ and nothing else""")
# Step 4: See what the cloud changes
print("\n what a managed platform changes, and what it does not:")
print(f" {'':<24}{'your laptop':<24}{'managed platform'}")
for label, local, cloud in (
("the algorithm", "scikit-learn", "SCIKIT-LEARN -- identical"),
("data source", "a local file", "s3:// or a feature store"),
("hardware", "what you own", "chosen per job, per hour"),
("training time", "hours on CPU", "minutes on GPU, if it helps"),
("experiment tracking", "a notebook cell", "logged automatically"),
("the artefact", "a file you might lose", "versioned in object storage"),
("deployment", "you build a server", "one API call"),
("cost", "sunk", "PER SECOND, and visible")):
print(f" {label:<24}{local:<24}{cloud}")
print(f""" THE FIRST ROW IS THE POINT. The cloud does not make your
model better; it makes training REPRODUCIBLE, deployment
ROUTINE and cost VISIBLE. A bad model trained on 8 GPUs is
still a bad model, and the {dummy_acc:.0%} baseline above is unmoved by
any amount of hardware""")
# Step 5: Price the instance
print(f"\n what this training job would cost (it took "
f"{train_seconds:.2f}s here):")
print(f" {'instance':<16}{'$/hour':>9}{'10 min job':>12}"
f"{'100 jobs':>11}")
for inst in ("t3.medium", "m5.xlarge", "c5.4xlarge", "p3.2xlarge",
"p4d.24xlarge"):
rate = f.EC2[inst]
job = rate / 6
print(f" {inst:<16}{rate:>9.4f}{job:>12.4f}{job * 100:>11.2f}")
cheap = f.EC2["m5.xlarge"] / 6
gpu = f.EC2["p4d.24xlarge"] / 6
assert gpu / cheap > 100
print(f""" the 8-GPU box costs {gpu / cheap:.0f}x the general-purpose one for the
same ten minutes. GPUs earn that on deep learning, where the
work is dense matrix multiplication. GRADIENT BOOSTING ON
TABULAR DATA DOES NOT USE THEM -- this model would run at the
same speed and 170x the price.
'Which instance?' is answered by the ALGORITHM, not by
ambition""")
# Step 6: Experiment 14: run AutoML
print("\n --- experiment 14: AutoML, actually run")
candidates = {
"LogisticRegression": Pipeline([
("s", StandardScaler()),
("c", LogisticRegression(max_iter=2000, random_state=SEED))]),
"DecisionTree(d=3)": DecisionTreeClassifier(max_depth=3,
random_state=SEED),
"DecisionTree(d=None)": DecisionTreeClassifier(random_state=SEED),
"RandomForest(100)": RandomForestClassifier(n_estimators=100,
random_state=SEED),
"GradientBoosting": GradientBoostingClassifier(random_state=SEED),
}
print(f" {len(candidates)} candidates, 5-fold CV on ROC AUC "
f"-- a real search")
board = []
total_fits = 0
t0 = time.perf_counter()
for name, est in candidates.items():
scores = cross_val_score(est, X_train, y_train, cv=5,
scoring="roc_auc")
total_fits += 5
board.append((name, scores.mean(), scores.std()))
search_seconds = time.perf_counter() - t0
board.sort(key=lambda r: -r[1])
print(f"\n {'rank':<6}{'model':<24}{'CV AUC':>9}{'std':>8}")
for i, (name, mean, sd) in enumerate(board, 1):
print(f" {i:<6}{name:<24}{mean:>9.4f}{sd:>8.4f}")
winner, best, best_sd = board[0]
second, second_score, second_sd = board[1]
print(f"\n {total_fits} model fits in {search_seconds:.1f}s")
assert len(board) == 5
gap = best - second_score
print(f""" THE LEADERBOARD IS THE WHOLE OF AUTOML. It fits many
models, cross-validates each, and ranks them. There is no
intelligence in it -- it is a SEARCH, and its value is that
it is exhaustive where you would have been lazy.
And look at the top two: {best:.4f} against {second_score:.4f}, a gap of
{gap:.4f} with standard deviations of {best_sd:.4f} and {second_sd:.4f}. THE
DIFFERENCE IS INSIDE THE NOISE. Declaring a winner here is
not supported by the data, and 'AutoML picked X' is not a
reason to prefer X""")
# Step 7: List what AutoML cannot do
print("\n what AutoML does NOT do:")
for what in (
"decide what the target variable should be",
"notice that your target leaks the answer",
"tell you the base rate matters more than the algorithm",
"know that last year's data no longer describes this year",
"choose a threshold that fits the business cost of an error",
"explain a prediction to a regulator",
"notice that the model is unfair to a protected group",
):
print(f" - {what}")
print(""" EVERY ONE OF THOSE IS THE ACTUAL JOB. AutoML automates
the part a competent person does in an afternoon and leaves
untouched the parts that take weeks and cause the failures.
Say that when asked to evaluate AutoML, and say it before
saying it is useful -- which it is""")
# Step 8: Price the search
print("\n what an AutoML search costs")
per_fit_here = search_seconds / total_fits
print(f" one fit on THIS dataset (1,200 rows): "
f"{per_fit_here:.3f} s -- too small to cost anything")
print(" so scale it: assume one fit takes 4 minutes, which is "
"ordinary")
print(f"\n {'search':<26}{'fits':>6}{'compute':>12}{'m5.xlarge':>12}")
MINUTES_PER_FIT = 4
for label, fits in (("this search, scaled", total_fits),
("a modest managed search", 250),
("a full AutoML run", 2000)):
hours = fits * MINUTES_PER_FIT / 60
cost = hours * f.EC2["m5.xlarge"]
print(f" {label:<26}{fits:>6}{hours:>10.1f} h"
f" ${cost:>9.2f}")
one = 1 * MINUTES_PER_FIT / 60 * f.EC2["m5.xlarge"]
full = 2000 * MINUTES_PER_FIT / 60 * f.EC2["m5.xlarge"]
assert full / one == 2000
print(f""" AutoML's compute is a straight MULTIPLE of one fit --
${full:,.2f} against ${one:.4f}, exactly {full / one:,.0f}x, because that is
all it is. And AutoML services charge a premium on top of
the compute.
So cut the search space with what you already know: on
tabular data, gradient boosting wins often enough that
starting there and stopping is frequently the better trade.
The leaderboard above makes the point -- it spent 5x the
compute to rank a model {gap:.4f} AUC above the one you would
have picked anyway, inside the noise""")
os.remove(MODEL_PATH)
return board
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 11_sagemaker_train.md:
NOT RUN HERE
11_sagemaker_train.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:
OUTPUT
Experiments 11 and 14 -- training and AutoML on a managed platform
dataset: 1,200 rows, 10 features, base rate 15.67% positive
--- experiment 11: the training job
model accuracy F1 AUC
DummyClassifier 0.8433 0.0000 0.5000
GradientBoosting 0.9467 0.8095 0.9029
quote the DUMMY FIRST, always. 94.67% sounds excellent
until you see that predicting 'never churns' scores 84.33%
-- the real gain is 10.3% points of accuracy, and it is the
F1 of 0.8095 against 0.0000 that shows the model found anything
at all. This is Course 12 A's argument, and it does not stop
being true because the model trained on somebody else's
computer
model artefact: 138,945 bytes, reloads and predicts identically
THAT FILE IS THE DELIVERABLE. A SageMaker training job
writes exactly this to s3://bucket/models/, and the deploy
step reads it back. Training and serving are separate
systems joined by one artefact in object storage -- which is
why the IAM role in experiment 10 needs s3:PutObject on
models/ and nothing else
what a managed platform changes, and what it does not:
your laptop managed platform
the algorithm scikit-learn SCIKIT-LEARN -- identical
data source a local file s3:// or a feature store
hardware what you own chosen per job, per hour
training time hours on CPU minutes on GPU, if it helps
experiment tracking a notebook cell logged automatically
the artefact a file you might lose versioned in object storage
deployment you build a server one API call
cost sunk PER SECOND, and visible
THE FIRST ROW IS THE POINT. The cloud does not make your
model better; it makes training REPRODUCIBLE, deployment
ROUTINE and cost VISIBLE. A bad model trained on 8 GPUs is
still a bad model, and the 84% baseline above is unmoved by
any amount of hardware
what this training job would cost (it took 0.27s here):
instance $/hour 10 min job 100 jobs
t3.medium 0.0416 0.0069 0.69
m5.xlarge 0.1920 0.0320 3.20
c5.4xlarge 0.6800 0.1133 11.33
p3.2xlarge 3.0600 0.5100 51.00
p4d.24xlarge 32.7726 5.4621 546.21
the 8-GPU box costs 171x the general-purpose one for the
same ten minutes. GPUs earn that on deep learning, where the
work is dense matrix multiplication. GRADIENT BOOSTING ON
TABULAR DATA DOES NOT USE THEM -- this model would run at the
same speed and 170x the price.
'Which instance?' is answered by the ALGORITHM, not by
ambition
--- experiment 14: AutoML, actually run
5 candidates, 5-fold CV on ROC AUC -- a real search
rank model CV AUC std
1 RandomForest(100) 0.9334 0.0210
2 GradientBoosting 0.9288 0.0196
3 DecisionTree(d=None) 0.8213 0.0425
4 LogisticRegression 0.8154 0.0606
5 DecisionTree(d=3) 0.8025 0.0364
25 model fits in 2.2s
THE LEADERBOARD IS THE WHOLE OF AUTOML. It fits many
models, cross-validates each, and ranks them. There is no
intelligence in it -- it is a SEARCH, and its value is that
it is exhaustive where you would have been lazy.
And look at the top two: 0.9334 against 0.9288, a gap of
0.0047 with standard deviations of 0.0210 and 0.0196. THE
DIFFERENCE IS INSIDE THE NOISE. Declaring a winner here is
not supported by the data, and 'AutoML picked X' is not a
reason to prefer X
what AutoML does NOT do:
- decide what the target variable should be
- notice that your target leaks the answer
- tell you the base rate matters more than the algorithm
- know that last year's data no longer describes this year
- choose a threshold that fits the business cost of an error
- explain a prediction to a regulator
- notice that the model is unfair to a protected group
EVERY ONE OF THOSE IS THE ACTUAL JOB. AutoML automates
the part a competent person does in an afternoon and leaves
untouched the parts that take weeks and cause the failures.
Say that when asked to evaluate AutoML, and say it before
saying it is useful -- which it is
what an AutoML search costs
one fit on THIS dataset (1,200 rows): 0.089 s -- too small to cost anything
so scale it: assume one fit takes 4 minutes, which is ordinary
search fits compute m5.xlarge
this search, scaled 25 1.7 h $ 0.32
a modest managed search 250 16.7 h $ 3.20
a full AutoML run 2000 133.3 h $ 25.60
AutoML's compute is a straight MULTIPLE of one fit --
$25.60 against $0.0128, exactly 2,000x, because that is
all it is. And AutoML services charge a premium on top of
the compute.
So cut the search space with what you already know: on
tabular data, gradient boosting wins often enough that
starting there and stopping is frequently the better trade.
The leaderboard above makes the point -- it spent 5x the
compute to rank a model 0.0047 AUC above the one you would
have picked anyway, inside the noise
The artefact is the deliverable. 138,945 bytes, written to disk, reloaded, and
predicting identically. A SageMaker training job writes exactly this to s3://bucket/models/,
and the deploy step reads it back. Training and serving are separate systems joined by one
file in object storage — which is why the IAM role in experiment 10 needs s3:PutObject on
models/ and nothing else.
The instance choice, priced:
| Instance | $/hour | 10-min job |
|---|---|---|
| m5.xlarge | 0.1920 | 0.0320 |
| p3.2xlarge (1 GPU) | 3.0600 | 0.5100 |
| p4d.24xlarge (8 GPU) | 32.7726 | 5.4621 |
171× for the same ten minutes — and gradient boosting on tabular data has no GPU code path. It would run at exactly the same speed.
NOTE
"Which instance?" is answered by the algorithm, not by ambition.
The three times the program prints — the training job, the 25 fits, and one fit — measure this machine at one moment and differ from run to run. Corrected: this page said one fit takes 0.127 s; that was one run's figure.
RESULT
Gradient boosting reaches 94.67% against the dummy's 84.33%, F1 0.8095 against 0; the artefact reloads and predicts identically.
Run a simple ETL job: extract, transform, and load into a cloud warehouse.
Build the job in Glue, load Redshift or BigQuery, and reconcile the counts.
The procedure, on the console and CLI, 12_etl_to_warehouse.md:
ETL AGAINST ELT
ELT won because warehouse compute got cheap and elastic. Landing raw data
means a transformation bug is fixed by re-running SQL rather than
re-extracting from a production database that may no longer hold the old rows
— exactly the DELETE problem Big Data Technologies found in Sqoop.
The procedure, on the console and CLI, 12_etl_to_warehouse.md:
# Experiment 12 -- a simple ETL job: extract, transform, load into a cloud warehouse
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`09_etl_warehouse.py`, which runs the whole pipeline into a real columnar warehouse**.
---
<!-- Step 1: Build a Glue job -->
## AWS Glue (the managed option)
```python
import sys
from awsglue.transforms import *
from awsglue.context import GlueContext
from pyspark.context import SparkContext
glueContext = GlueContext(SparkContext.getOrCreate())
src = glueContext.create_dynamic_frame.from_catalog(
database="retail", table_name="orders")
mapped = ApplyMapping.apply(frame=src, mappings=[
("order_id", "long", "order_id", "long"),
("region", "string", "region", "string"),
("qty", "long", "qty", "int"),
("list_price", "double", "revenue_unit", "double"),
])
clean = DropNullFields.apply(frame=mapped)
glueContext.write_dynamic_frame.from_options(
frame=clean, connection_type="s3",
connection_options={"path": "s3://retail-lake/curated/",
"partitionKeys": ["quarter"]},
format="parquet")
```
**`partitionKeys` and `format="parquet"` are the two lines that matter**, and
they are Course 12 B's partition pruning and columnar storage arriving in the
cloud unchanged.
<!-- Step 2: Run the crawler -->
## The crawler, and its trap
```bash
aws glue start-crawler --name retail-crawler
```
A crawler infers the schema and registers the table in the Data Catalog.
**It infers types from a sample**, so a column that is integer in the sample
and text later becomes a runtime failure. For anything that matters, define
the table explicitly.
<!-- Step 3: Load Redshift -->
## Loading Redshift
```sql
COPY sales FROM 's3://retail-lake/curated/'
IAM_ROLE 'arn:aws:iam::<acct>:role/RedshiftCopyRole'
FORMAT AS PARQUET;
```
**`COPY` is the only sensible way to load Redshift.** A loop of `INSERT`
statements is orders of magnitude slower, because a columnar store is built
for bulk loads and each `INSERT` writes a nearly-empty block.
```sql
ANALYZE sales; -- statistics, or the optimiser guesses
VACUUM sales; -- reclaim space and re-sort after deletes
```
<!-- Step 4: Load BigQuery -->
## Loading BigQuery
```bash
bq load --source_format=PARQUET \
--time_partitioning_field=order_date \
--clustering_fields=region,category \
retail.sales gs://retail-lake/curated/*.parquet
```
**Partition on the column you filter by; cluster on the columns you filter
and group by after that.** Partitioning removes whole days from the scan;
clustering sorts within a partition so the block statistics can skip more.
Both directly reduce **bytes scanned**, which is directly the bill.
<!-- Step 5: Reconcile the counts -->
## The reconciliation nobody does and everybody should
```sql
SELECT COUNT(*) AS rows, SUM(revenue) AS revenue FROM sales;
```
Compare against the source. **Row count and the sum of a money column** —
the same two checks Course 12 B used on the Sqoop import. Rows alone will not
catch a truncated numeric type.
<!-- Step 6: Orchestrate it -->
## Orchestration
| Tool | Shape |
|---|---|
| Step Functions | a state machine, JSON, AWS-native |
| **Airflow (MWAA/Composer)** | a Python DAG — the industry standard |
| Glue Workflows | Glue jobs only |
| EventBridge + Lambda | event-driven, for small steps |
**Retries and idempotency are the whole design problem.** A pipeline that
cannot be re-run safely will eventually be re-run anyway, at 3 a.m., by
someone who does not know what it does.
The procedure, on the console and CLI, 12_etl_to_warehouse.md:
NOT RUN HERE
12_etl_to_warehouse.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model for this experiment, 09_etl_warehouse.py, is shown in full under Experiment 9, with what it printed.
RESULT
The same pipeline as Experiment 9, on managed services; ELT won because warehouse compute got cheap.
Use CloudWatch or Stackdriver to monitor endpoints, set alarms and configure auto-scaling.
Simulate a day of traffic, autoscale it, tune the thresholds, and choose what to alarm on.
The procedure, on the console and CLI, 13_monitoring.md:
The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:
THE DAY
A day of traffic: peak 1,000 req/s, trough 164; instances serve 150 req/s; the group is 2–12.
| Strategy | Instance-hours | Dropped |
|---|---|---|
| fixed at peak (7) | 168 | 0 |
| autoscaled, 70%/40%, cooldown 1 | 129 | 1,014 |
The procedure, on the console and CLI, 13_monitoring.md:
# Experiment 13 -- use CloudWatch/Stackdriver to monitor endpoints, set alarms and auto-scale
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`13_monitoring_autoscale.py`, which runs the control loop and measures what tuning costs**.
---
<!-- Step 1: Create an alarm -->
## An alarm
```bash
aws cloudwatch put-metric-alarm \
--alarm-name endpoint-p99-latency \
--namespace AWS/SageMaker \
--metric-name ModelLatency \
--dimensions Name=EndpointName,Value=churn-endpoint \
Name=VariantName,Value=AllTraffic \
--extended-statistic p99 \
--period 60 --evaluation-periods 3 \
--threshold 500000 --comparison-operator GreaterThanThreshold \
--treat-missing-data notBreaching \
--alarm-actions arn:aws:sns:us-east-1:<acct>:oncall
```
**`--extended-statistic p99`, not `--statistic Average`.** The runnable half
shows twenty latencies where one request took 900 ms: the mean is 85 ms and
would never fire an alarm, while p99 is 737 ms.
**`ModelLatency` is in MICROSECONDS.** 500000 is half a second. Getting the
unit wrong is how an alarm is set 1,000x too high and never fires.
**`--treat-missing-data`** decides what "no data" means. `notBreaching` is
right for a bursty endpoint; `breaching` is right when silence itself is the
failure.
<!-- Step 2: Set the billing alarm first -->
## The billing alarm, which comes first
```bash
aws cloudwatch put-metric-alarm --alarm-name monthly-spend \
--namespace AWS/Billing --metric-name EstimatedCharges \
--dimensions Name=Currency,Value=USD \
--statistic Maximum --period 21600 --evaluation-periods 1 \
--threshold 50 --comparison-operator GreaterThanThreshold \
--alarm-actions arn:aws:sns:us-east-1:<acct>:oncall
```
**Billing metrics exist only in `us-east-1`** regardless of where you work,
and they lag by about six hours. Set this before anything else in the course.
<!-- Step 3: Autoscale an endpoint -->
## Auto-scaling a SageMaker endpoint
```bash
aws application-autoscaling register-scalable-target \
--service-namespace sagemaker \
--resource-id endpoint/churn-endpoint/variant/AllTraffic \
--scalable-dimension sagemaker:variant:DesiredInstanceCount \
--min-capacity 1 --max-capacity 8
aws application-autoscaling put-scaling-policy \
--policy-name track-invocations --policy-type TargetTrackingScaling \
--service-namespace sagemaker \
--resource-id endpoint/churn-endpoint/variant/AllTraffic \
--scalable-dimension sagemaker:variant:DesiredInstanceCount \
--target-tracking-scaling-policy-configuration '{
"TargetValue": 750.0,
"PredefinedMetricSpecification":
{"PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"},
"ScaleInCooldown": 300, "ScaleOutCooldown": 60}'
```
**`ScaleOutCooldown` short and `ScaleInCooldown` long.** Scale out eagerly
because being under capacity drops requests; scale in reluctantly because
scaling back out costs boot time. The runnable half measures what happens
when you get that backwards.
<!-- Step 4: Read what the runnable half shows -->
## What the runnable half shows, and it is not flattering
- **Autoscaling dropped 1,014 requests where fixed capacity dropped none**,
because the group is always sized for the previous observation.
- **The most aggressive configuration cost MORE than fixed capacity** —
188 instance-hours against 168 — by overshooting.
**"Autoscaling saves money" is a claim about a tuned autoscaler.** Say that,
and give the numbers.
<!-- Step 5: Choose the six metrics -->
## The six metrics worth alarming on
| Metric | Alarm when | The trap |
|---|---|---|
| `ModelLatency` p99 | > 500 ms for 3 min | the mean hides it |
| `Invocation5XXErrors` | > 0 for 1 min | these are **yours** |
| `Invocation4XXErrors` | > 1% of requests | a rate, never a count |
| `CPUUtilization` | > 70% for 5 min | I/O-bound apps never reach it |
| `EstimatedCharges` | > your budget | lags ~6 h, `us-east-1` only |
| **`Invocations` == 0** | for 1 hour | **a dead endpoint still bills** |
**The last row is the one people miss.** An endpoint serving nothing looks
perfect on every performance metric and costs exactly the same as a busy one.
The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:
"""Experiment 13 -- CloudWatch/Stackdriver: monitor an endpoint, set alarms,
and configure auto-scaling rules.
The console steps are in `13_monitoring.md`, marked as not run here.
What runs here is the CONTROL LOOP, which is the part that actually behaves
in ways people do not expect: scaling lags demand, aggressive thresholds
oscillate, cooldowns trade responsiveness for stability, and an alarm on an
average hides the tail. Every one of those is measured below rather than
asserted.
"""
import fixtures as f
CAPACITY_PER_INSTANCE = 150 # requests/sec one instance can serve
MIN_INSTANCES = 2
MAX_INSTANCES = 12
def simulate(traffic, scale_out_at, scale_in_at, cooldown,
start=MIN_INSTANCES, step=1):
"""A target-tracking autoscaler, one tick per hour.
The key realism: a scaling decision made at tick t only takes effect at
tick t+1. Real instances take minutes to boot, so the group is ALWAYS
sized for the PREVIOUS observation.
"""
instances = start
cool = 0
history = []
for demand in traffic:
capacity = instances * CAPACITY_PER_INSTANCE
util = demand / capacity
served = min(demand, capacity)
dropped = demand - served
history.append({"demand": demand, "instances": instances,
"util": util, "dropped": dropped})
if cool > 0:
cool -= 1
elif util > scale_out_at and instances < MAX_INSTANCES:
instances = min(MAX_INSTANCES, instances + step)
cool = cooldown
elif util < scale_in_at and instances > MIN_INSTANCES:
instances = max(MIN_INSTANCES, instances - step)
cool = cooldown
return history
def summarise(history):
hours = len(history)
return {
"instance_hours": sum(h["instances"] for h in history),
"dropped": sum(h["dropped"] for h in history),
"peak_instances": max(h["instances"] for h in history),
"hours_over_90": sum(1 for h in history if h["util"] > 0.90),
"mean_util": sum(h["util"] for h in history) / hours,
"changes": sum(1 for a, b in zip(history, history[1:])
if a["instances"] != b["instances"]),
}
def percentile(values, p):
s = sorted(values)
k = (len(s) - 1) * p / 100
lo, hi = int(k), min(int(k) + 1, len(s) - 1)
return s[lo] + (s[hi] - s[lo]) * (k - lo)
def main():
print(" Experiment 13 -- monitoring, alarms and auto-scaling")
# Step 1: Read a day of traffic
traffic = f.daily_traffic()
print(f"\n a day of traffic: peak {max(traffic)} req/s, "
f"trough {min(traffic)} req/s, {sum(traffic):,} req/s-hours")
print(f" one instance serves {CAPACITY_PER_INSTANCE} req/s; "
f"group is {MIN_INSTANCES}..{MAX_INSTANCES}")
# Step 2: Fix the capacity at the peak
print("\n OPTION 1 -- fixed capacity, sized for peak:")
need = -(-max(traffic) // CAPACITY_PER_INSTANCE)
fixed = [{"demand": d, "instances": need,
"util": d / (need * CAPACITY_PER_INSTANCE),
"dropped": 0} for d in traffic]
fs = summarise(fixed)
print(f" {need} instances all day = {fs['instance_hours']} "
f"instance-hours, 0 dropped")
print(f" mean utilisation {fs['mean_util'] * 100:.1f}%")
assert fs["dropped"] == 0
print(""" nothing is dropped and most of the fleet is idle most of
the day. That is the pre-cloud bargain: you buy the peak and
pay for it at 3 a.m.""")
# Step 3: Autoscale
print("\n OPTION 2 -- target tracking, scale out above 70%, "
"in below 40%, cooldown 1:")
auto = simulate(traffic, 0.70, 0.40, cooldown=1)
a = summarise(auto)
print(f" {'hour':>5}{'demand':>8}{'inst':>6}{'util':>8}{'dropped':>9}")
for h, row in enumerate(auto):
flag = " <-- shortfall" if row["dropped"] else ""
print(f" {h:>5}{row['demand']:>8}{row['instances']:>6}"
f"{row['util'] * 100:>7.0f}%{row['dropped']:>9}{flag}")
saved = fs["instance_hours"] - a["instance_hours"]
pct = 100 * saved / fs["instance_hours"]
print(f"\n instance-hours {a['instance_hours']} against "
f"{fs['instance_hours']} fixed -- {pct:.0f}% fewer")
print(f" requests dropped: {a['dropped']:,}")
print(f" scaling changes : {a['changes']}")
assert a["instance_hours"] < fs["instance_hours"]
assert a["dropped"] > 0
# Step 4: Read it honestly
worst = max(auto, key=lambda r: r["dropped"])
hour = auto.index(worst)
print(f"""
AUTOSCALING DROPPED {a['dropped']:,} REQUESTS AND FIXED CAPACITY DROPPED
NONE. The worst hour is {hour}, where demand jumped to {worst['demand']}
against {worst['instances']} instances -- the group was sized for the
PREVIOUS hour, because a scaling decision takes effect one
tick late.
Autoscaling does not track demand. It CHASES demand, and it
is always one observation behind. That lag is the cost of the
{pct:.0f}% saving, and pretending otherwise is how a launch goes
badly""")
# Step 5: Tune the thresholds
print("\n the same day at different thresholds:")
print(f" {'out/in':<12}{'cool':>5}{'inst-hrs':>10}{'dropped':>9}"
f"{'changes':>9}{'mean util':>11}")
configs = [(0.70, 0.40, 1), (0.50, 0.30, 1), (0.85, 0.60, 1),
(0.70, 0.40, 3), (0.50, 0.30, 0)]
results = {}
for out, inn, cd in configs:
r = summarise(simulate(traffic, out, inn, cd))
results[(out, inn, cd)] = r
print(f" {f'{out:.0%}/{inn:.0%}':<12}{cd:>5}"
f"{r['instance_hours']:>10}{r['dropped']:>9}"
f"{r['changes']:>9}{r['mean_util'] * 100:>10.0f}%")
aggressive = results[(0.50, 0.30, 1)]
lazy = results[(0.85, 0.60, 1)]
assert aggressive["dropped"] < lazy["dropped"]
assert aggressive["instance_hours"] > lazy["instance_hours"]
print(f""" THE TABLE IS A TRADE-OFF CURVE, NOT A LEADERBOARD.
Scaling out at 50% drops {aggressive['dropped']:,} requests for {aggressive['instance_hours']} instance-hours;
scaling out at 85% drops {lazy['dropped']:,} for {lazy['instance_hours']}. You are choosing
between spare capacity and dropped requests, and only a
business can say which is worse""")
no_cool = results[(0.50, 0.30, 0)]
assert no_cool["dropped"] == 0
assert no_cool["instance_hours"] > fs["instance_hours"]
print(f"""
AND READ THE LAST ROW AGAINST FIXED CAPACITY. Scaling out at
50% with NO cooldown drops nothing -- and costs {no_cool['instance_hours']}
instance-hours against fixed capacity's {fs['instance_hours']}.
AUTOSCALING MADE IT MORE EXPENSIVE. Chase demand hard enough
and the group overshoots on the way up and lingers on the way
down, so you buy more than the peak. 'Autoscaling saves
money' is a claim about a TUNED autoscaler, not about
autoscaling""")
long_cool = results[(0.70, 0.40, 3)]
short_cool = results[(0.70, 0.40, 1)]
print(f"""
and the cooldown: 3 ticks gives {long_cool['changes']} scaling changes against
{short_cool['changes']}, at {long_cool['dropped'] - short_cool['dropped']:+,} dropped requests. A long cooldown
stops FLAPPING -- scaling out and back in repeatedly around a
threshold, which costs boot time and stabilises nothing""")
# Step 6: Choose what to alarm on
print("\n alarms: what to measure, and the trap in each")
lat = [40, 42, 41, 45, 43, 40, 44, 42, 41, 43,
41, 42, 40, 44, 43, 41, 42, 45, 40, 900]
mean = sum(lat) / len(lat)
p50, p95, p99 = (percentile(lat, p) for p in (50, 95, 99))
print(f" 20 request latencies, one of them 900 ms:")
print(f" mean {mean:.1f} ms p50 {p50:.1f} ms "
f"p95 {p95:.1f} ms p99 {p99:.1f} ms")
assert mean < 100 and p99 > 400
print(f""" AN ALARM ON THE MEAN ({mean:.0f} ms) NEVER FIRES. One request
in twenty took 900 ms and the average absorbed it. Alarm on
p95 or p99, because the tail is where users live -- and note
that 5% of requests is a lot of users""")
print(f"\n {'metric':<22}{'alarm when':<26}{'the trap'}")
for m, when, trap in (
("CPUUtilization", "> 70% for 5 min", "an I/O-bound app never hits it"),
("ModelLatency p99", "> 500 ms for 3 min", "the mean hides it"),
("Invocation4XXErrors", "> 1% of requests", "a rate, never a raw count"),
("Invocation5XXErrors", "> 0 for 1 min", "these are YOUR fault"),
("EstimatedCharges", "> your monthly budget", "billing metrics lag ~6 h"),
("no invocations at all", "== 0 for 1 hour", "a dead endpoint still bills")):
print(f" {m:<22}{when:<26}{trap}")
print(""" THE LAST ROW IS THE ONE PEOPLE MISS. An endpoint serving
nothing looks perfect on every performance metric and costs
the same as a busy one. Alarm on ABSENCE of traffic, and
alarm on spend -- those two catch the failures that
monitoring dashboards are blind to""")
# Step 7: Price the day
print("\n what the day cost, at m5.large on-demand:")
rate = f.EC2["m5.large"]
for label, hours in (("fixed at peak", fs["instance_hours"]),
("autoscaled", a["instance_hours"])):
print(f" {label:<18}{hours:>5} instance-hours "
f"${hours * rate:>7.2f}/day ${hours * rate * 30:>8.2f}/month")
fixed_month = fs["instance_hours"] * rate * 30
auto_month = a["instance_hours"] * rate * 30
spot_month = auto_month * (1 - f.SPOT_DISCOUNT)
res_month = fixed_month * (1 - f.RESERVED_DISCOUNT)
print(f"\n autoscaled on SPOT (-{f.SPOT_DISCOUNT:.0%}, interruptible): "
f"${spot_month:,.2f}/month")
print(f" fixed on RESERVED (-{f.RESERVED_DISCOUNT:.0%}, 1-yr commit): "
f"${res_month:,.2f}/month")
assert spot_month < auto_month < fixed_month
print(f""" the real answer is usually BOTH: a reserved baseline for
the floor you always need, autoscaled on-demand or spot for
the peak. Here that is {MIN_INSTANCES} reserved instances plus the rest
elastic -- and it beats either pure strategy, which is why
every cost-optimisation review starts by asking what your
floor is""")
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 13_monitoring.md:
NOT RUN HERE
13_monitoring.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:
OUTPUT
Experiment 13 -- monitoring, alarms and auto-scaling
a day of traffic: peak 1000 req/s, trough 164 req/s, 13,996 req/s-hours
one instance serves 150 req/s; group is 2..12
OPTION 1 -- fixed capacity, sized for peak:
7 instances all day = 168 instance-hours, 0 dropped
mean utilisation 55.5%
nothing is dropped and most of the fleet is idle most of
the day. That is the pre-cloud bargain: you buy the peak and
pay for it at 3 a.m.
OPTION 2 -- target tracking, scale out above 70%, in below 40%, cooldown 1:
hour demand inst util dropped
0 208 2 69% 0
1 182 2 61% 0
2 164 2 55% 0
3 164 2 55% 0
4 173 2 58% 0
5 226 2 75% 0
6 366 3 81% 0
7 604 3 134% 154 <-- shortfall
8 824 4 137% 224 <-- shortfall
9 1000 4 167% 400 <-- shortfall
10 930 5 124% 180 <-- shortfall
11 806 5 107% 56 <-- shortfall
12 736 6 82% 0
13 701 6 78% 0
14 666 7 63% 0
15 692 7 66% 0
16 771 7 73% 0
17 894 8 74% 0
18 965 8 80% 0
19 912 9 68% 0
20 754 9 56% 0
21 578 9 43% 0
22 402 9 30% 0
23 278 8 23% 0
instance-hours 129 against 168 fixed -- 23% fewer
requests dropped: 1,014
scaling changes : 8
AUTOSCALING DROPPED 1,014 REQUESTS AND FIXED CAPACITY DROPPED
NONE. The worst hour is 9, where demand jumped to 1000
against 4 instances -- the group was sized for the
PREVIOUS hour, because a scaling decision takes effect one
tick late.
Autoscaling does not track demand. It CHASES demand, and it
is always one observation behind. That lag is the cost of the
23% saving, and pretending otherwise is how a launch goes
badly
the same day at different thresholds:
out/in cool inst-hrs dropped changes mean util
70%/40% 1 129 1014 8 77%
50%/30% 1 158 358 10 62%
85%/60% 1 114 1380 7 87%
70%/40% 3 96 2093 4 98%
50%/30% 0 188 0 11 52%
THE TABLE IS A TRADE-OFF CURVE, NOT A LEADERBOARD.
Scaling out at 50% drops 358 requests for 158 instance-hours;
scaling out at 85% drops 1,380 for 114. You are choosing
between spare capacity and dropped requests, and only a
business can say which is worse
AND READ THE LAST ROW AGAINST FIXED CAPACITY. Scaling out at
50% with NO cooldown drops nothing -- and costs 188
instance-hours against fixed capacity's 168.
AUTOSCALING MADE IT MORE EXPENSIVE. Chase demand hard enough
and the group overshoots on the way up and lingers on the way
down, so you buy more than the peak. 'Autoscaling saves
money' is a claim about a TUNED autoscaler, not about
autoscaling
and the cooldown: 3 ticks gives 4 scaling changes against
8, at +1,079 dropped requests. A long cooldown
stops FLAPPING -- scaling out and back in repeatedly around a
threshold, which costs boot time and stabilises nothing
alarms: what to measure, and the trap in each
20 request latencies, one of them 900 ms:
mean 85.0 ms p50 42.0 ms p95 87.8 ms p99 737.5 ms
AN ALARM ON THE MEAN (85 ms) NEVER FIRES. One request
in twenty took 900 ms and the average absorbed it. Alarm on
p95 or p99, because the tail is where users live -- and note
that 5% of requests is a lot of users
metric alarm when the trap
CPUUtilization > 70% for 5 min an I/O-bound app never hits it
ModelLatency p99 > 500 ms for 3 min the mean hides it
Invocation4XXErrors > 1% of requests a rate, never a raw count
Invocation5XXErrors > 0 for 1 min these are YOUR fault
EstimatedCharges > your monthly budget billing metrics lag ~6 h
no invocations at all == 0 for 1 hour a dead endpoint still bills
THE LAST ROW IS THE ONE PEOPLE MISS. An endpoint serving
nothing looks perfect on every performance metric and costs
the same as a busy one. Alarm on ABSENCE of traffic, and
alarm on spend -- those two catch the failures that
monitoring dashboards are blind to
what the day cost, at m5.large on-demand:
fixed at peak 168 instance-hours $ 16.13/day $ 483.84/month
autoscaled 129 instance-hours $ 12.38/day $ 371.52/month
autoscaled on SPOT (-70%, interruptible): $111.46/month
fixed on RESERVED (-40%, 1-yr commit): $290.30/month
the real answer is usually BOTH: a reserved baseline for
the floor you always need, autoscaled on-demand or spot for
the peak. Here that is 2 reserved instances plus the rest
elastic -- and it beats either pure strategy, which is why
every cost-optimisation review starts by asking what your
floor is
AUTOSCALING DROPPED 1,014 REQUESTS AND FIXED CAPACITY DROPPED NONE
The worst hour is hour 9: demand jumped to 1,000 against 4 instances, because the group was sized for the previous hour.
NOTE
Autoscaling does not track demand. It CHASES demand, and it is always one observation behind.
That lag is the cost of the 23% saving.
The tuning curve:
| Out/in | Cool | Inst-hrs | Dropped | Changes |
|---|---|---|---|---|
| 70%/40% | 1 | 129 | 1,014 | 8 |
| 50%/30% | 1 | 158 | 358 | 10 |
| 85%/60% | 1 | 114 | 1,380 | 7 |
| 70%/40% | 3 | 96 | 2,093 | 4 |
| 50%/30% | 0 | 188 | 0 | 11 |
This is a trade-off curve, not a leaderboard. You are choosing between spare capacity and dropped requests, and only a business can say which is worse.
AND READ THE LAST ROW AGAINST FIXED CAPACITY
188 instance-hours against 168. Scaling out at 50% with no cooldown drops nothing — and costs more than simply buying the peak.
NOTE
Autoscaling made it MORE expensive. Chase demand hard enough and the group overshoots on the way up and lingers on the way down.
"Autoscaling saves money" is a claim about a tuned autoscaler.
The cooldown row: 3 ticks gives 4 scaling changes instead of 8, at +1,079 dropped requests. A long cooldown stops flapping — which costs boot time and stabilises nothing. Scale out eagerly, scale in reluctantly.
Alarm on the tail:
20 latencies, one of them 900 ms:
mean 85.0 ms p50 42.0 ms p95 87.8 ms p99 737.5 ms
An alarm on the mean never fires. Alarm on p95 or p99 — the tail is where users live, and 5% of requests is a lot of users.
THE METRIC NOBODY SETS
| Metric | Alarm when | The trap |
|---|---|---|
ModelLatency p99 |
> 500 ms | the mean hides it; the unit is microseconds |
Invocation5XXErrors |
> 0 | these are yours |
Invocation4XXErrors |
> 1% | a rate, never a count |
EstimatedCharges |
> budget | lags ~6 h, us-east-1 only |
Invocations == 0 |
for 1 hour | a dead endpoint still bills |
The last row is the one people miss. An endpoint serving nothing looks perfect on every performance metric and costs the same as a busy one.
And what the day cost:
| Per month | |
|---|---|
| fixed at peak, on-demand | $483.84 |
| autoscaled, on-demand | $371.52 |
| fixed, reserved (−40%) | $290.30 |
| autoscaled, spot (−70%) | $111.46 |
The real answer is usually both: a reserved baseline for the floor, spot or on-demand for the peak.
RESULT
Autoscaling saved 23% of instance-hours and dropped 1,014 requests that fixed capacity kept; tuned hard enough, it cost more than fixed capacity.
Use a cloud AutoML service for a prediction task.
Run a model search, read its leaderboard, and see what it does not do.
The procedure, on the console and CLI, 14_automl.md:
AUTOML, ACTUALLY RUN
Five candidates, 5-fold CV, 25 real fits:
| Rank | Model | CV AUC | std |
|---|---|---|---|
| 1 | RandomForest(100) | 0.9334 | 0.0210 |
| 2 | GradientBoosting | 0.9288 | 0.0196 |
| 3 | DecisionTree(depth=None) | 0.8213 | 0.0425 |
| 4 | LogisticRegression | 0.8154 | 0.0606 |
| 5 | DecisionTree(depth=3) | 0.8025 | 0.0364 |
The leaderboard is the whole of AutoML. Fit many models, cross-validate, rank. There is no intelligence in it — it is a search.
The procedure, on the console and CLI, 14_automl.md:
# Experiment 14 -- use cloud AutoML services for a dataset prediction task
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`11_train_and_automl.py`, which runs a real 5-model, 5-fold search and reports the leaderboard**.
---
<!-- Step 1: Run SageMaker Autopilot -->
## SageMaker Autopilot
```python
from sagemaker.automl.automl import AutoML
automl = AutoML(
role=role,
target_attribute_name="churned",
output_path=f"s3://{bucket}/autopilot/",
problem_type="BinaryClassification",
job_objective={"MetricName": "AUC"},
max_candidates=20,
max_runtime_per_training_job_in_seconds=600,
)
automl.fit(inputs=f"s3://{bucket}/train/train.csv")
automl.describe_auto_ml_job()["BestCandidate"]
```
**`max_candidates` and `max_runtime_per_training_job_in_seconds` are the
budget**, and they are not optional. Without them the job explores until it
is satisfied, and it bills the whole time.
<!-- Step 2: Or Vertex AI AutoML -->
## Vertex AI AutoML
```bash
gcloud ai custom-jobs create --region=us-central1 ...
# or, tabular:
gcloud beta ai models list --region=us-central1
```
Vertex bills AutoML in **node-hours** with a documented minimum. Read the
minimum before starting — a small dataset does not produce a small bill.
<!-- Step 3: Or Azure Automated ML -->
## Azure Automated ML
```python
from azure.ai.ml import automl
job = automl.classification(
training_data=train, target_column_name="churned",
primary_metric="AUC_weighted",
experiment_timeout_minutes=30, # THE BUDGET
enable_early_termination=True,
)
```
<!-- Step 4: See what they do -->
## What these actually do
**They fit many models, cross-validate each, and rank them.** The runnable
half does exactly this with five candidates and 5-fold CV, and prints the
leaderboard. There is no intelligence in it — it is a **search**, and its
value is that it is exhaustive where a human would be lazy.
**And read the top of that leaderboard carefully.** In the run here the top
two models differ by 0.0047 AUC with standard deviations of 0.0210 and
0.0196. **The difference is inside the noise**, and "AutoML picked X" is not
a reason to prefer X.
<!-- Step 5: And what they do not -->
## What AutoML does not do
- decide what the target variable should be
- notice that your target **leaks** the answer
- tell you the base rate matters more than the algorithm
- know that last year's data no longer describes this year
- choose a threshold that fits the business cost of an error
- explain a prediction to a regulator
- notice the model is unfair to a protected group
**Every one of those is the actual job.** AutoML automates the afternoon and
leaves the weeks untouched.
<!-- Step 6: Read the explainability report -->
## The explainability report
Autopilot generates a candidate-definition notebook and a data-exploration
notebook. **Read them** — they are the best thing about the product, because
they show you the feature engineering it chose, which is the part you would
otherwise never see and could not defend.
The procedure, on the console and CLI, 14_automl.md:
NOT RUN HERE
14_automl.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
AND THE TOP TWO ARE INSIDE THE NOISE
0.9334 against 0.9288 is a gap of 0.0047, with standard deviations of 0.0210 and 0.0196. Declaring a winner is not supported by the data, and "AutoML picked X" is not a reason to prefer X.
What the search costs. One fit here takes a fraction of a second — too small to cost anything. Scale to a realistic four minutes per fit:
| Search | Fits | Compute | m5.xlarge |
|---|---|---|---|
| this search | 25 | 1.7 h | $0.32 |
| a modest managed search | 250 | 16.7 h | $3.20 |
| a full AutoML run | 2,000 | 133.3 h | $25.60 |
A straight multiple of one fit — exactly 2,000× — because that is all it is. And managed services charge a premium on top.
WHAT AUTOML DOES NOT DO
decide the target · notice leakage · tell you the base rate matters more · know last year's data no longer applies · choose a threshold that fits the business cost · explain a prediction · notice unfairness
Every one of those is the actual job. AutoML automates the afternoon and leaves the weeks untouched.
The Python model for this experiment, 11_train_and_automl.py, is shown in full under Experiment 11, with what it printed.
RESULT
Five candidates, 25 real fits: random forest leads gradient boosting by 0.0047, inside the noise.
Deploy a trained ML model as a REST API endpoint.
Serve the model over HTTP, check health, predictions and errors, and measure latency and batching.
The procedure, on the console and CLI, 15_deploy.md:
The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:
THE EQUALITY THAT IS THE DEPLOYMENT TEST
artefact loaded from disk: 138,945 bytes
GET /ping -> 200 {'status': 'healthy', 'model_loaded': True}
POST /invocations (3 rows) -> 200
predictions : [0, 0, 0]
probabilities : [0.015383, 0.008643, 0.014492]
POST /invocations (300 rows) -> 200, accuracy 0.9467
The endpoint's answers are identical to calling the model in-process. Serving must not change predictions — and a preprocessing step that lives in your notebook rather than in the pipeline is exactly how it does.
The procedure, on the console and CLI, 15_deploy.md:
# Experiment 15 -- deploy a trained ML model as a REST API endpoint
## NOT EXECUTED
**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.
So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.
The runnable half is **`15_deploy_endpoint.py`, which starts a REAL HTTP server, serves a REAL model and calls it**.
---
<!-- Step 1: Deploy -->
## Deploy
```python
predictor = estimator.deploy(
initial_instance_count=1,
instance_type="ml.m5.large",
endpoint_name="churn-endpoint",
)
predictor.predict([[1.2, 0.4, ...]])
```
Or from the CLI, in the three steps the SDK hides:
```bash
aws sagemaker create-model --model-name churn-model \
--primary-container Image=<ecr-uri>,ModelDataUrl=s3://bucket/models/model.tar.gz \
--execution-role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole
aws sagemaker create-endpoint-config --endpoint-config-name churn-config \
--production-variants VariantName=AllTraffic,ModelName=churn-model,\
InitialInstanceCount=1,InstanceType=ml.m5.large,InitialVariantWeight=1
aws sagemaker create-endpoint --endpoint-name churn-endpoint \
--endpoint-config-name churn-config
```
**Model, endpoint config, endpoint — three objects, not one.** That is what
makes blue/green possible: create a second config and update the endpoint,
and traffic shifts without downtime.
<!-- Step 2: Invoke it -->
## Invoke it
```bash
aws sagemaker-runtime invoke-endpoint \
--endpoint-name churn-endpoint \
--content-type application/json \
--body '{"instances": [[1.2, 0.4, 3.3, ...]]}' \
/dev/stdout
```
**The request is IAM-signed.** There is no API key to leak, and access is
governed by the same policy evaluation as everything else.
<!-- Step 3: Follow the container contract -->
## The container contract
Your container must answer two routes:
| Route | Must |
|---|---|
| `GET /ping` | return 200 quickly, **without running the model** |
| `POST /invocations` | run inference |
**A health check that does real inference marks the container unhealthy
whenever the model is merely slow — and the platform then kills a container
that was working.** The runnable half implements both routes and says so.
<!-- Step 4: Choose real-time, serverless or batch -->
## Real-time, serverless or batch
| | Real-time | Serverless | Batch transform |
|---|---|---|---|
| Latency | ms | ms, **after a cold start** | minutes to hours |
| Billed | **per hour, always** | per request | per job |
| Idle cost | **the full instance** | **zero** | zero |
| Good for | steady traffic | spiky or occasional | scoring a whole file |
**An `ml.m5.large` endpoint is about $70/month whether or not anything calls
it.** If traffic is occasional, serverless inference costs a fraction; if you
are scoring a file, batch transform is the right tool and an endpoint is the
expensive way to do arithmetic — the runnable half measures a **37x**
difference between one batched request and 100 single ones.
<!-- Step 5: Delete it -->
## Then delete it
```bash
aws sagemaker delete-endpoint --endpoint-name churn-endpoint
aws sagemaker delete-endpoint-config --endpoint-config-name churn-config
aws sagemaker delete-model --model-name churn-model
```
**Deleting the endpoint is a step in the experiment, not an afterthought.**
Every "surprise AWS bill" story is a resource nobody switched off.
<!-- Step 6: Monitor it -->
## Monitoring the deployed model
Data drift is the failure that has no error message: the endpoint keeps
returning 200 and the predictions quietly stop being right.
```python
from sagemaker.model_monitor import DefaultModelMonitor
monitor = DefaultModelMonitor(role=role, instance_type="ml.m5.xlarge")
monitor.suggest_baseline(baseline_dataset=f"s3://{bucket}/train/train.csv")
monitor.create_monitoring_schedule(endpoint_input=predictor.endpoint_name,
schedule_cron_expression="cron(0 * ? * * *)")
```
**Baseline the training distribution, then compare production inputs against
it hourly.** A drift alarm is the only thing that catches a model that has
stopped working while every infrastructure metric stays green.
The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:
"""Experiment 15 -- deploy a trained ML model as a REST API endpoint.
`15_deploy.md` carries the SageMaker deploy call and the console steps, NOT
EXECUTED -- there is no cloud account.
But the ENDPOINT ITSELF RUNS HERE. A real HTTP server starts on localhost,
serves a real scikit-learn model over a real JSON API, is called with real
requests, and is shut down. The contract, the error handling, the health
check and the latency measurements are genuine -- only the hosting is not.
That is the honest split: SageMaker gives you a container, a load balancer,
autoscaling and an IAM-signed URL. What it serves is this.
"""
import json
import os
import tempfile
import threading
import time
import urllib.error
import urllib.request
from http.server import BaseHTTPRequestHandler, HTTPServer
import joblib
import numpy as np
import fixtures as f
from sklearn.datasets import make_classification
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
SEED = 42
N_FEATURES = 10
HOST = "127.0.0.1"
PORT = 0 # the OS picks a free port; the real one is read back
_MODEL = None
_STATE = {"invocations": 0, "errors_4xx": 0, "errors_5xx": 0,
"latencies": []}
class Handler(BaseHTTPRequestHandler):
"""The two routes every model endpoint must have, and no more."""
def log_message(self, *args):
pass # keep the suite output clean
def _send(self, code, payload):
body = json.dumps(payload).encode()
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(body)))
self.end_headers()
self.wfile.write(body)
def do_GET(self):
if self.path == "/ping":
# SageMaker calls THIS to decide whether the container is alive.
# It must not run the model: a health check that does real work
# takes the endpoint down when the model is merely slow.
self._send(200, {"status": "healthy",
"model_loaded": _MODEL is not None})
elif self.path == "/metrics":
self._send(200, dict(_STATE, latencies=len(_STATE["latencies"])))
else:
_STATE["errors_4xx"] += 1
self._send(404, {"error": "not found"})
def do_POST(self):
if self.path != "/invocations":
_STATE["errors_4xx"] += 1
self._send(404, {"error": "not found"})
return
started = time.perf_counter()
try:
length = int(self.headers.get("Content-Length", 0))
payload = json.loads(self.rfile.read(length) or b"{}")
rows = payload.get("instances")
if not isinstance(rows, list) or not rows:
raise ValueError("body must be {'instances': [[...], ...]}")
arr = np.asarray(rows, dtype=float)
if arr.ndim != 2 or arr.shape[1] != N_FEATURES:
raise ValueError(
f"each instance needs {N_FEATURES} features, "
f"got shape {list(arr.shape)}")
except (ValueError, TypeError, json.JSONDecodeError) as exc:
# A BAD REQUEST IS A 4XX, NOT A 5XX. Getting this wrong makes
# your error alarm fire for other people's mistakes.
_STATE["errors_4xx"] += 1
self._send(400, {"error": str(exc)})
return
try:
proba = _MODEL.predict_proba(arr)[:, 1]
preds = (proba >= 0.5).astype(int)
except Exception as exc: # pragma: no cover
_STATE["errors_5xx"] += 1
self._send(500, {"error": "inference failed"})
return
_STATE["invocations"] += len(rows)
_STATE["latencies"].append((time.perf_counter() - started) * 1000)
self._send(200, {"predictions": preds.tolist(),
"probabilities": [round(p, 6) for p in proba]})
def build_model():
X, y = make_classification(
n_samples=1200, n_features=N_FEATURES, n_informative=5,
n_redundant=2, weights=[0.85, 0.15], flip_y=0.02,
class_sep=1.1, random_state=SEED)
Xtr, Xte, ytr, yte = train_test_split(
X, y, test_size=0.25, stratify=y, random_state=SEED)
pipe = Pipeline([("scale", StandardScaler()),
("clf", GradientBoostingClassifier(random_state=SEED))])
pipe.fit(Xtr, ytr)
return pipe, Xte, yte
_PORT = None
def call(path, payload=None, method="GET"):
url = f"http://{HOST}:{_PORT}{path}"
data = json.dumps(payload).encode() if payload is not None else None
req = urllib.request.Request(
url, data=data, method=method,
headers={"Content-Type": "application/json"} if data else {})
try:
with urllib.request.urlopen(req, timeout=10) as resp:
return resp.status, json.loads(resp.read())
except urllib.error.HTTPError as exc:
return exc.code, json.loads(exc.read())
def percentile(values, p):
s = sorted(values)
k = (len(s) - 1) * p / 100
lo, hi = int(k), min(int(k) + 1, len(s) - 1)
return s[lo] + (s[hi] - s[lo]) * (k - lo)
def main():
global _MODEL, _PORT
print(" Experiment 15 -- a model deployed as a REST endpoint, "
"actually served")
# Step 1: Train the model, and save it
model, X_test, y_test = build_model()
path = os.path.join(tempfile.gettempdir(), "cloud13b_endpoint.joblib")
joblib.dump(model, path)
_MODEL = joblib.load(path)
print(f"\n artefact loaded from disk: {os.path.getsize(path):,} bytes")
# Step 2: Serve it
HTTPServer.allow_reuse_address = True
server = HTTPServer((HOST, PORT), Handler)
global _PORT
_PORT = server.server_address[1]
thread = threading.Thread(target=server.serve_forever, daemon=True)
thread.start()
# [Changed: this printed the port, which the system picks afresh on every run.]
print(f" endpoint listening on http://{HOST}, on a port the system chose "
f"(a REAL HTTP server)")
try:
# Step 3: Check its health
code, body = call("/ping")
print(f"\n GET /ping -> {code} {body}")
assert code == 200 and body["model_loaded"] is True
print(""" /ping answers WITHOUT running the model. A health check
that does real inference marks the container unhealthy
whenever the model is merely slow, and the platform then
kills a container that was working -- a self-inflicted
outage, and a classic one""")
# Step 4: Ask for a prediction
sample = X_test[:3].tolist()
code, body = call("/invocations", {"instances": sample}, "POST")
print(f"\n POST /invocations with 3 rows -> {code}")
print(f" predictions : {body['predictions']}")
print(f" probabilities : {body['probabilities']}")
assert code == 200 and len(body["predictions"]) == 3
local = model.predict(X_test[:3]).tolist()
assert body["predictions"] == local
print(""" the endpoint's answers are IDENTICAL to calling the model
in-process. That equality is the deployment test worth
writing: serving must not change predictions, and a
preprocessing step that lives in your notebook rather than
in the pipeline is exactly how it does""")
# Step 5: Send a batch
code, body = call("/invocations",
{"instances": X_test.tolist()}, "POST")
assert code == 200
preds = np.array(body["predictions"])
acc = (preds == y_test).mean()
print(f"\n POST /invocations with all {len(X_test)} rows -> {code}, "
f"accuracy {acc:.4f}")
assert acc > 0.90
# Step 6: Send bad requests
print("\n error handling, which is most of a real endpoint:")
cases = [
("wrong feature count", {"instances": [[1.0, 2.0]]}, "POST",
"/invocations"),
("not a list", {"instances": "hello"}, "POST", "/invocations"),
("empty body", {}, "POST", "/invocations"),
("wrong route", {"instances": sample}, "POST", "/predict"),
("wrong route, GET", None, "GET", "/predict"),
]
print(f" {'case':<24}{'status':>8} message")
for label, payload, method, route in cases:
code, body = call(route, payload, method)
msg = body.get("error", "")[:46]
print(f" {label:<24}{code:>8} {msg}")
assert 400 <= code < 500, "a client mistake must not be a 5xx"
print(""" EVERY ONE IS A 4XX, NOT A 5XX, and that distinction is
operational rather than pedantic: 5xx means YOUR service is
broken and should page someone. If malformed client input
returns 500, your error alarm fires for other people's bugs
and you stop trusting it""")
# Step 7: Measure the latency
print("\n latency over 200 single-row requests:")
_STATE["latencies"].clear()
for i in range(200):
code, _ = call("/invocations",
{"instances": [X_test[i % len(X_test)].tolist()]},
"POST")
assert code == 200
lat = _STATE["latencies"]
p50, p95, p99 = (percentile(lat, p) for p in (50, 95, 99))
mean = sum(lat) / len(lat)
print(f" mean {mean:.3f} ms p50 {p50:.3f} ms "
f"p95 {p95:.3f} ms p99 {p99:.3f} ms")
assert p99 >= p50
print(f""" p99 is {p99 / p50:.1f}x p50 on an idle laptop serving one model.
On a shared endpoint under load that ratio grows, which is
why the alarm in experiment 13 is on p99 and not the mean.
These are SERVER-SIDE numbers; a client also pays network
time, and the user's experience is the sum""")
# Step 8: Batch, against one at a time
print("\n one request of 100 rows against 100 requests of one row:")
_STATE["latencies"].clear()
t0 = time.perf_counter()
call("/invocations", {"instances": X_test[:100].tolist()}, "POST")
batched = (time.perf_counter() - t0) * 1000
t0 = time.perf_counter()
for i in range(100):
call("/invocations", {"instances": [X_test[i].tolist()]}, "POST")
singly = (time.perf_counter() - t0) * 1000
print(f" batched : {batched:8.2f} ms total")
print(f" one by one: {singly:8.2f} ms total "
f"({singly / batched:.0f}x)")
assert singly > batched
print(f""" {singly / batched:.0f}x, and none of it is the model -- it is per-request
overhead: HTTP, JSON parsing, and a NumPy call whose fixed
cost is paid 100 times instead of once.
This is why batch transform exists alongside real-time
endpoints. If you are scoring a file of a million rows,
calling an endpoint a million times is the expensive way to
do arithmetic""")
# Step 9: Read the metrics
code, metrics = call("/metrics")
print(f"\n GET /metrics -> {metrics['invocations']:,} invocations, "
f"{metrics['errors_4xx']} 4xx, {metrics['errors_5xx']} 5xx")
assert metrics["errors_5xx"] == 0
finally:
server.shutdown()
server.server_close()
os.remove(path)
print("\n endpoint shut down and artefact removed.")
# Step 10: Compare with a managed endpoint
print("\n what SageMaker adds that this server does not have:")
print(f" {'':<26}{'this script':<22}{'a managed endpoint'}")
for label, here, cloud in (
("TLS", "no", "yes, terminated for you"),
("authentication", "NONE -- anyone", "IAM-signed requests"),
("load balancing", "one process", "across instances and AZs"),
("autoscaling", "no", "on InvocationsPerInstance"),
("blue/green deploy", "no", "traffic shifted gradually"),
("metrics", "the dict above", "CloudWatch, automatically"),
("cost", "electricity", "PER HOUR, until deleted")):
print(f" {label:<26}{here:<22}{cloud}")
hourly = f.EC2["m5.large"]
print(f"\n an ml.m5.large endpoint: ${hourly:.4f}/hour "
f"= ${hourly * f.HOURS_PER_MONTH:,.2f}/month, called or not")
print(""" THE LAST ROW AGAIN. A training job stops; an endpoint does
not. Deleting the endpoint is a step in the experiment, not
an afterthought -- and if traffic is occasional, a serverless
endpoint or batch transform costs a fraction of it""")
if __name__ == "__main__":
main()
The procedure, on the console and CLI, 15_deploy.md:
NOT RUN HERE
15_deploy.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.
The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:
OUTPUT
Experiment 15 -- a model deployed as a REST endpoint, actually served
artefact loaded from disk: 138,945 bytes
endpoint listening on http://127.0.0.1, on a port the system chose (a REAL HTTP server)
GET /ping -> 200 {'status': 'healthy', 'model_loaded': True}
/ping answers WITHOUT running the model. A health check
that does real inference marks the container unhealthy
whenever the model is merely slow, and the platform then
kills a container that was working -- a self-inflicted
outage, and a classic one
POST /invocations with 3 rows -> 200
predictions : [0, 0, 0]
probabilities : [0.015383, 0.008643, 0.014492]
the endpoint's answers are IDENTICAL to calling the model
in-process. That equality is the deployment test worth
writing: serving must not change predictions, and a
preprocessing step that lives in your notebook rather than
in the pipeline is exactly how it does
POST /invocations with all 300 rows -> 200, accuracy 0.9467
error handling, which is most of a real endpoint:
case status message
wrong feature count 400 each instance needs 10 features, got shape [1,
not a list 400 body must be {'instances': [[...], ...]}
empty body 400 body must be {'instances': [[...], ...]}
wrong route 404 not found
wrong route, GET 404 not found
EVERY ONE IS A 4XX, NOT A 5XX, and that distinction is
operational rather than pedantic: 5xx means YOUR service is
broken and should page someone. If malformed client input
returns 500, your error alarm fires for other people's bugs
and you stop trusting it
latency over 200 single-row requests:
mean 0.735 ms p50 0.709 ms p95 0.972 ms p99 1.184 ms
p99 is 1.7x p50 on an idle laptop serving one model.
On a shared endpoint under load that ratio grows, which is
why the alarm in experiment 13 is on p99 and not the mean.
These are SERVER-SIDE numbers; a client also pays network
time, and the user's experience is the sum
one request of 100 rows against 100 requests of one row:
batched : 2.74 ms total
one by one: 155.69 ms total (57x)
57x, and none of it is the model -- it is per-request
overhead: HTTP, JSON parsing, and a NumPy call whose fixed
cost is paid 100 times instead of once.
This is why batch transform exists alongside real-time
endpoints. If you are scoring a file of a million rows,
calling an endpoint a million times is the expensive way to
do arithmetic
GET /metrics -> 703 invocations, 5 4xx, 0 5xx
endpoint shut down and artefact removed.
what SageMaker adds that this server does not have:
this script a managed endpoint
TLS no yes, terminated for you
authentication NONE -- anyone IAM-signed requests
load balancing one process across instances and AZs
autoscaling no on InvocationsPerInstance
blue/green deploy no traffic shifted gradually
metrics the dict above CloudWatch, automatically
cost electricity PER HOUR, until deleted
an ml.m5.large endpoint: $0.0960/hour = $70.08/month, called or not
THE LAST ROW AGAIN. A training job stops; an endpoint does
not. Deleting the endpoint is a step in the experiment, not
an afterthought -- and if traffic is occasional, a serverless
endpoint or batch transform costs a fraction of it
/PING MUST NOT RUN THE MODEL
A health check that does real inference marks the container unhealthy whenever the model is merely slow — and the platform then kills a container that was working. A self-inflicted outage, and a classic one.
Error handling, which is most of a real endpoint:
| Case | Status |
|---|---|
| wrong feature count | 400 |
| body not a list | 400 |
| empty body | 400 |
| wrong route (POST and GET) | 404 |
Every one is a 4xx, not a 5xx, and the distinction is operational rather than pedantic: 5xx means your service is broken and should page someone. If malformed client input returns 500, your error alarm fires for other people's bugs and you stop trusting it.
Latency, over 200 real requests, and the batching result, are the timed lines above: they measure this machine at one moment and differ from run to run. What does not change is their shape — p99 above p50 even on an idle machine serving one model, and one request of 100 rows far faster than 100 requests of one row, none of the difference being the model: it is per-request overhead, HTTP, JSON parsing, and a NumPy call whose fixed cost is paid 100 times instead of once. Corrected: this page gave "p99 is 1.8× p50" and "37×" as if fixed; they were one run's figures, and three runs made here gave batching factors of 47×, 57× and 68×. Under load the p99 ratio grows — which is why experiment 13's alarm is on p99.
NOTE
If you are scoring a million rows, calling an endpoint a million times is the expensive way to do arithmetic.
What SageMaker adds that this server does not have:
| This script | A managed endpoint | |
|---|---|---|
| TLS | no | terminated for you |
| Authentication | NONE — anyone | IAM-signed requests |
| Load balancing | one process | across instances and AZs |
| Autoscaling | no | on InvocationsPerInstance |
| Blue/green | no | traffic shifted gradually |
| Cost | electricity | $70/month, called or not |
Deleting the endpoint is a step in the experiment, not an afterthought.
Changed: the program printed the port it listened on, which the system picks afresh on every run; it now says so.
RESULT
The endpoint's predictions equal the model's in-process; bad input gets a 4xx, never a 5xx; batching beat one row per request many times over.
| Script | Experiments | Real? |
|---|---|---|
01_vm_and_hosting.py |
1, 2, 7 | a real web server; overcommit modelled |
03_iam_and_account.py |
3, 10 | the real IAM algorithm |
04_storage.py |
4, 5, 6 | real key semantics, real arithmetic |
09_etl_warehouse.py |
8, 9, 12 | real SQLite → real DuckDB |
11_train_and_automl.py |
11, 14 | a real model, a real 25-fit search |
13_monitoring_autoscale.py |
13 | a real control loop |
15_deploy_endpoint.py |
15 | a real HTTP endpoint, called over TCP |
Plus the audit: 14 Markdown files, every one carrying NOT EXECUTED, each naming the service it needs.
Two hours on a console, one experiment number, then a viva.
What costs marks:
*:* in an IAM policy, and calling it "it works now"s3 mv as a renameLIMIT 10 makes a BigQuery query cheapWhat earns them:
The three IAM rules, applied. Explicit deny; else allow; else deny — and the demonstration that adding S3 admin changes nothing.
"Prefix-scoped Deny beats bucket-scoped Allow." One sentence that explains a data lake's raw zone.
The two storage-class ratios: 23× on storage, 8.7× all-in. And Standard winning outright at two retrievals a month.
"1 TB out costs what 3.9 months of storage costs." Egress, data gravity and lock-in in one figure.
The 500× BigQuery difference, and naming it as Big Data Technologies' column projection and partition pruning saving money instead of time.
The 254 TB break-even. Serverless against provisioned as a calculation.
"The AutoML gap is inside the noise." 0.0047 against standard deviations of 0.02.
"Autoscaling made it more expensive." 188 instance-hours against 168 — reporting the result that contradicts the slogan.
Batching against one row per request, argued with the number your run measured.
These experiments are console procedures rather than programs, so each one is written out as a page.
The same experiments, one page each, so a program can be reached by what it does rather than by its number.