Skip to the content
On this page
  1. Read this before you read anything else
  2. Experiment 1 — Create a virtual machine
  3. Experiment 2 — Host a page on the server
  4. Experiment 3 — Create and configure a cloud account
  5. Experiment 4 — Create and manage storage buckets
  6. Experiment 5 — Block storage
  7. Experiment 6 — File storage
  8. Experiment 7 — The notebook environment
  9. Experiment 8 — Cloud-hosted databases
  10. Experiment 9 — A batch ETL pipeline
  11. Experiment 10 — A SageMaker notebook with an IAM role
  12. Experiment 11 — Train a model on a managed platform
  13. Experiment 12 — An ETL job into a cloud warehouse
  14. Experiment 13 — Monitoring, alarms and auto-scaling
  15. Experiment 14 — AutoML
  16. Experiment 15 — Deploy the model as a REST endpoint
  17. What the runner asserts
  18. Lab examination
  19. Written-out instructions
  20. Each program, on its own page

15 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.

Code lives in labs/course-13b-cloud/.

Read this before you read anything else

NOTE

There is no cloud account for this repository, and none will be created.

Signing up for AWS, Azure or GCP requires a payment card and accepts a billing relationship. That is not a thing a study repository should do on anyone's behalf, so no provider was ever contacted and no claim in these notes about a provider's behaviour was demonstrated here.

Half Files Status
The console and CLI steps 14 Markdown files NOT EXECUTED, at the top of every one: each experiment's 4. Programme is its procedure, and 5. Execution and Results says so in a box
The verification 7 programs Executed and asserted by tools/data-science/run_cloud_labs.py; what each printed is under the experiment it is named for
pip install -r tools/requirements.txt
python3 tools/data-science/run_cloud_labs.py

AND YET A SURPRISING AMOUNT REALLY RUNS

Most of what this course teaches is not proprietary:

Runs for real What it is
IAM policy evaluation the actual algorithm, in iam.py
Object-store semantics prefixes, no directories, copy-plus-delete, versioning
All the pricing arithmetic storage classes, egress, per-TB, per-node-hour
Hypervisor overcommit and the point at which it fails
A real web server serving a real page over TCP, fetched back
A real ETL pipeline SQLite → transform → DuckDB, with an audit trail
An autoscaling control loop measured, including where it loses
A real model and a real AutoML search scikit-learn, 25 real fits
A real REST endpoint serving that model, called over the network

Nothing is claimed that was not executed. Every .md file names the service it needs and the runnable half that verifies its logic, and the runner asserts the marker is still present.

Three of the programs time something — a training job, a model search, an endpoint's latency. A timing measures the machine at a moment, so those lines differ from run to run, and capture_lab_outputs.py --check sets them aside; every other line must repeat exactly.

The cross-course check

Experiments 8, 9 and 12 use Business Intelligence Tools' star schema, imported not copied. ₹10,360 for South is now produced by Business Intelligence Tools' DAX, Big Data Technologies' Hive, Big Data Technologies' Spark and this course's DuckDB — four engines, nine facts, and verify_all.sh fails if any of them drifts.


Experiment 1 — Create a virtual machine

1. Question

Create a virtual machine in VMware Workstation.

2. Aim

Run the new-VM wizard, make the three choices that matter, and see where memory overcommit fails.

3. Steps

The procedure, on the console and CLI, 01_create_vm.md:

  1. Run the wizard.
  2. Make the three choices that matter.
  3. Finish the install.

The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:

  1. Experiment 1: allocate the guests, and overcommit.
  2. Experiment 2: serve a page.
  3. Experiment 7: run the notebook's cells.

THE THREE WIZARD CHOICES THAT GET PEOPLE

4. Programme

The procedure, on the console and CLI, 01_create_vm.md:

# Experiment 1 -- create a virtual machine in VMware Workstation

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`01_vm_and_hosting.py`, which models overcommit and measures where it breaks**.

---

<!-- Step 1: Run the wizard -->
## The wizard, and what each choice actually means

| Step | Choice | Why it matters |
|---|---|---|
| New Virtual Machine | **Custom (advanced)** | Typical hides the disk and network options you need |
| Hardware compatibility | Workstation 17.x | older only if you must move the VM to an older host |
| Guest OS install | **Install later** | attach the ISO after, so the wizard does not run an unattended install |
| Guest OS | Linux → Ubuntu 64-bit | picks sensible defaults for the virtual chipset |
| Processors | 2 cores | more than the host has *physical* cores makes it slower, not faster |
| Memory | 4096 MB | see the overcommit note below |
| Network | **NAT** | the guest shares the host's IP; Bridged gives it one on your LAN |
| Disk controller | NVMe (or SCSI) | IDE is slow and there is no reason for it |
| Disk | 40 GB, **split**, **not** pre-allocated | see the disk note below |

<!-- Step 2: Make the three choices that matter -->
## The three choices that get people

**1. Memory.** The host must keep enough for itself. Giving a 16 GB laptop's
VM 12 GB leaves 4 GB for Windows, the hypervisor and your browser, and the
whole machine swaps. **Type 2 hypervisors do not balloon aggressively**, so
the overcommit that works in a datacentre does not work here.

**2. Disk: "Allocate all disk space now" vs "split into multiple files".**
Pre-allocating writes 40 GB immediately and is faster afterwards. Not
pre-allocating grows on demand and is what you want on a laptop. Splitting
into 2 GB files matters only for filesystems that cannot hold a 40 GB file
(FAT32) — and for copying the VM to a USB stick.

**3. NAT vs Bridged vs Host-only.**

| Mode | The guest gets | Reachable from the LAN? |
|---|---|---|
| **NAT** | a private IP behind the host | **no**, unless you forward a port |
| **Bridged** | an IP from your router's DHCP | **yes** |
| Host-only | a private IP, no internet | no |

**Choose NAT and then wonder why nobody can reach your web server** — that is
experiment 2's most common failure, and the fix is either Bridged or a port
forward in `Edit → Virtual Network Editor`.

<!-- Step 3: Finish the install -->
## After the install

```bash
sudo apt update && sudo apt install -y open-vm-tools open-vm-tools-desktop
# ^ shared clipboard, drag-and-drop, correct screen resolution
ip addr show          # note the guest's IP -- you need it for experiment 2
free -h ; nproc ; df -h
```

**Take a snapshot before you install anything else.** `VM → Snapshot → Take
Snapshot`. A snapshot is the only reason experimenting in a VM is safe, and
it is the feature that has no cheap equivalent on physical hardware.

## The link forward

Every EC2 instance, Azure VM and GCE instance is a guest on a **type 1**
hypervisor — ESXi, Hyper-V, KVM or AWS's Nitro. **The cloud is this
experiment at rack scale**, with the wizard replaced by an API call. Knowing
what the wizard was choosing is what makes an instance type comprehensible.

The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:

"""Experiments 1, 2 and 7 -- a virtual machine, a web server on it, and a
notebook environment.

VMware Workstation is not installed here, so `01_create_vm.md` carries the
wizard steps, marked as not run here. But the two things those experiments actually
teach DO run:

  * virtualization is RESOURCE MULTIPLEXING, and the interesting behaviour is
    overcommit -- modelled and measured below.
  * experiment 2 hosts a page on a server. THIS SCRIPT REALLY DOES THAT,
    with Python's own HTTP server standing in for Apache, and fetches the
    page back over TCP to prove it.
  * experiment 7 runs a notebook. Papermill and Jupyter are not installed,
    so the script executes the same cells directly and asserts the outputs,
    which is what a notebook test does anyway.
"""
import http.server
import json
import os
import socketserver
import tempfile
import threading
import urllib.error
import urllib.request

import fixtures as f

HOST = "127.0.0.1"
PORT = 0          # let the OS pick a free port -- see experiment_2


# ------------------------------------------------------------ experiment 1

def allocate(host_ram_gb, host_vcpu, vms, ballooning=True):
    """Place VMs on a host and report what the hypervisor actually does.

    Two facts drive everything:
      * vCPUs are TIME-SLICED, so you can allocate far more than you have.
      * RAM is not, at least not for free -- overcommitted memory is backed
        by ballooning, page sharing and finally SWAP, which is a cliff.
    """
    ram_alloc = sum(v["ram"] for v in vms)
    cpu_alloc = sum(v["vcpu"] for v in vms)
    ram_used = sum(v["ram"] * v["active"] for v in vms)
    reclaimed = (ram_alloc - ram_used) if ballooning else 0
    pressure = max(0.0, ram_used - host_ram_gb)
    return {
        "ram_allocated": ram_alloc,
        "ram_ratio": ram_alloc / host_ram_gb,
        "cpu_ratio": cpu_alloc / host_vcpu,
        "ram_actually_touched": ram_used,
        "reclaimed_by_ballooning": reclaimed,
        "swapping_gb": pressure,
    }


def experiment_1():
    print("\n    --- experiment 1: the virtual machine")
    print(f"      {'':<22}{'type 1 (bare metal)':<26}{'type 2 (hosted)'}")
    for label, t1, t2 in (
            ("runs on", "the hardware directly", "on top of an OS"),
            ("examples", "ESXi, Hyper-V, KVM, Xen", "VMware Workstation, VirtualBox"),
            ("overhead", "a few percent", "noticeably more"),
            ("used for", "datacentres, THE CLOUD", "a laptop, this experiment"),
            ("boots", "instead of an OS", "as an application")):
        print(f"      {label:<22}{t1:<26}{t2}")
    print("""         every EC2 instance, every Azure VM and every GCE instance
         is a guest on a TYPE 1 hypervisor. The whole cloud is this
         experiment, at rack scale -- which is why it is experiment 1""")

    host_ram, host_cpu = 32, 8
    vms = [
        {"name": "web-1",  "ram": 8,  "vcpu": 4, "active": 0.35},
        {"name": "web-2",  "ram": 8,  "vcpu": 4, "active": 0.30},
        {"name": "db-1",   "ram": 16, "vcpu": 4, "active": 0.90},
        {"name": "batch",  "ram": 16, "vcpu": 8, "active": 0.20},
    ]
    print(f"\n      a {host_ram} GB / {host_cpu} vCPU host, four guests:")
    print(f"      {'vm':<10}{'RAM':>6}{'vCPU':>6}{'active':>9}")
    for v in vms:
        print(f"      {v['name']:<10}{v['ram']:>5} G{v['vcpu']:>6}"
              f"{v['active']:>8.0%}")

    r = allocate(host_ram, host_cpu, vms)
    print(f"\n      allocated RAM  : {r['ram_allocated']} GB on a {host_ram} GB host "
          f"({r['ram_ratio']:.2f}x)")
    print(f"      allocated vCPU : {sum(v['vcpu'] for v in vms)} on {host_cpu} "
          f"({r['cpu_ratio']:.2f}x)")
    print(f"      RAM actually touched     : {r['ram_actually_touched']:.1f} GB")
    print(f"      reclaimed by ballooning  : {r['reclaimed_by_ballooning']:.1f} GB")
    print(f"      swapping                 : {r['swapping_gb']:.1f} GB")
    assert r["ram_ratio"] > 1 and r["cpu_ratio"] > 1
    assert r["swapping_gb"] == 0
    print(f"""         {r['ram_allocated']} GB allocated on a {host_ram} GB host and
         {sum(v['vcpu'] for v in vms)} vCPUs on {host_cpu}, and nothing is swapping -- because
         the guests only TOUCH {r['ram_actually_touched']:.1f} GB.
         Overcommit works on the same bet an airline makes, and it is
         why a cloud provider can sell more capacity than it owns""")

    print("\n      now the batch job wakes up (20% -> 95% active):")
    busy = [dict(v, active=0.95 if v["name"] == "batch" else v["active"])
            for v in vms]
    r2 = allocate(host_ram, host_cpu, busy)
    print(f"      RAM actually touched : {r2['ram_actually_touched']:.1f} GB")
    print(f"      swapping             : {r2['swapping_gb']:.1f} GB")
    assert r2["swapping_gb"] > 0
    print(f"""         {r2['swapping_gb']:.1f} GB OVER, AND NOW EVERY GUEST IS SLOW -- not just
         the batch job. Memory overcommit fails as a CLIFF, and it
         fails for the neighbours: this is the 'noisy neighbour'
         problem, and it is why cloud instance types quote DEDICATED
         memory and only burstable CPU.
         CPU overcommit degrades gracefully because time-slicing
         shares; RAM does not, because a page is either resident or
         it is not""")


# ------------------------------------------------------------ experiment 2

PAGE = """<!doctype html>
<title>Sales dashboard</title>
<h1>Retail sales</h1>
<table>
<tr><th>Region</th><th>Revenue</th></tr>
{rows}
</table>
<p>Total: {total}</p>
"""


def experiment_2():
    print("\n    --- experiment 2: host a page on the server (this RUNS)")
    doc_root = tempfile.mkdtemp(prefix="cloud13b_www_")

    by_region = (f.SALES_DF.groupby("region")["revenue"].sum()
                 .sort_values(ascending=False))
    rows = "\n".join(f"<tr><td>{k}</td><td>{v:,.0f}</td></tr>"
                     for k, v in by_region.items())
    html = PAGE.format(rows=rows, total=f"{f.total_revenue():,.0f}")
    with open(os.path.join(doc_root, "index.html"), "w") as fh:
        fh.write(html)
    with open(os.path.join(doc_root, "data.json"), "w") as fh:
        json.dump({k: float(v) for k, v in by_region.items()}, fh)

    class Quiet(http.server.SimpleHTTPRequestHandler):
        def __init__(self, *a, **kw):
            super().__init__(*a, directory=doc_root, **kw)

        def log_message(self, *a):
            pass

    class Reusable(socketserver.TCPServer):
        allow_reuse_address = True            # or a re-run hits TIME_WAIT

    with Reusable((HOST, PORT), Quiet) as httpd:
        port = httpd.server_address[1]
        t = threading.Thread(target=httpd.serve_forever, daemon=True)
        t.start()
        # [Changed: these printed the temporary folder's name and the port, which
        # the system picks afresh on every run.]
        print("      document root : a new temporary folder, cloud13b_www_...")
        print(f"      serving       : http://{HOST}, on a port the system chose  (a REAL server)")

        with urllib.request.urlopen(f"http://{HOST}:{port}/") as resp:
            served = resp.read().decode()
            ctype = resp.headers["Content-Type"]
            status = resp.status
        print(f"      GET /         -> {status}, {ctype}, "
              f"{len(served)} bytes")
        assert status == 200 and "text/html" in ctype
        assert "Retail sales" in served and "10,360" in served

        with urllib.request.urlopen(f"http://{HOST}:{port}/data.json") as resp:
            data = json.loads(resp.read())
            jtype = resp.headers["Content-Type"]
        print(f"      GET /data.json-> 200, {jtype}, {data}")
        assert data["South"] == 10360.0 and jtype == "application/json"

        try:
            urllib.request.urlopen(f"http://{HOST}:{port}/missing.html")
            raise AssertionError("should have 404ed")
        except urllib.error.HTTPError as exc:
            print(f"      GET /missing  -> {exc.code}")
            assert exc.code == 404

        httpd.shutdown()

    os.remove(os.path.join(doc_root, "index.html"))
    os.remove(os.path.join(doc_root, "data.json"))
    os.rmdir(doc_root)
    print("""         a page was written to a document root, served over TCP,
         fetched back, and its CONTENT-TYPE checked. That is the whole
         of experiment 2; Apache under XAMPP adds virtual hosts,
         .htaccess, PHP and TLS, and the shape is identical.
         Note the Content-Type header. A browser renders index.html
         because the server SAID text/html -- get that wrong and the
         browser downloads your page instead of showing it, which is
         the commonest 'my site is broken' on a fresh VM""")

    print(f"\n      {'concern':<24}{'on your VM':<26}{'managed (S3/App Service)'}")
    for c, vm, mg in (
            ("who patches Apache", "YOU, monthly", "the provider"),
            ("TLS certificate", "certbot, renewals", "issued and rotated"),
            ("scaling", "a bigger VM", "automatic"),
            ("a static site costs", "a VM, hourly", "cents per GB stored"),
            ("you control", "everything", "very little")):
        print(f"      {c:<24}{vm:<26}{mg}")
    print("""         a STATIC site on a VM is the clearest case of paying for
         a general-purpose computer to do something an object store
         does for cents. Hosting index.html on S3 + CloudFront costs
         less than the VM's first hour""")


# ------------------------------------------------------------ experiment 7

def experiment_7():
    print("\n    --- experiment 7: the notebook environment")
    cells = [
        ("import pandas as pd; import fixtures as f",
         lambda ns: ns.update({"df": f.SALES_DF}) or "ok"),
        ("df.shape", lambda ns: ns["df"].shape),
        ("df.groupby('region')['revenue'].sum().to_dict()",
         lambda ns: {k: float(v) for k, v in
                     ns["df"].groupby("region")["revenue"].sum().items()}),
        ("df['revenue'].sum()", lambda ns: float(ns["df"]["revenue"].sum())),
    ]
    ns, outputs = {}, []
    for src, fn in cells:
        out = fn(ns)
        outputs.append(out)
        shown = str(out)
        print(f"      In  [{len(outputs)}]: {src}")
        print(f"      Out [{len(outputs)}]: "
              f"{shown[:60]}{'...' if len(shown) > 60 else ''}")
    assert outputs[1] == (9, 19)
    assert outputs[2]["South"] == 10360.0
    assert outputs[3] == f.total_revenue()
    print("""         four cells, executed in order, every output asserted.
         That is what a notebook TEST looks like -- papermill or
         nbconvert --execute do exactly this in CI, and a notebook
         nobody executes in CI is a notebook that has already
         drifted""")

    print(f"\n      {'':<24}{'Colab':<24}{'notebook on a cloud VM'}")
    for label, colab, vm in (
            ("costs", "free tier, then paid", "the INSTANCE, hourly"),
            ("data access", "upload, or mount Drive", "IAM role, no keys"),
            ("stops when", "idle ~90 min", "NEVER -- you stop it"),
            ("state on stop", "LOST", "kept on the EBS volume"),
            ("GPU", "when available", "the one you pay for"),
            ("private data", "a policy question", "inside your VPC")):
        print(f"      {label:<24}{colab:<24}{vm}")
    idle = f.EC2["m5.xlarge"] * f.HOURS_PER_MONTH
    print(f"\n      an m5.xlarge notebook left running: ${idle:,.2f}/month")
    assert idle > 100
    print("""         'STOPS WHEN: NEVER' is the row that costs money. Colab
         disconnecting is an annoyance; a cloud notebook not
         disconnecting is a bill. Set an idle-shutdown lifecycle
         policy on day one -- SageMaker supports one, and it is the
         single most useful thing you can configure""")


def main():
    print("  Experiments 1, 2 and 7 -- VM, web server and notebook")
    # Step 1: Experiment 1: allocate the guests, and overcommit
    experiment_1()
    # Step 2: Experiment 2: serve a page
    experiment_2()
    # Step 3: Experiment 7: run the notebook's cells
    experiment_7()


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 01_create_vm.md:

NOT RUN HERE

01_create_vm.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 01_vm_and_hosting.py, for experiments 1, 2 and 7:

OUTPUT

  Experiments 1, 2 and 7 -- VM, web server and notebook

    --- experiment 1: the virtual machine
                            type 1 (bare metal)       type 2 (hosted)
      runs on               the hardware directly     on top of an OS
      examples              ESXi, Hyper-V, KVM, Xen   VMware Workstation, VirtualBox
      overhead              a few percent             noticeably more
      used for              datacentres, THE CLOUD    a laptop, this experiment
      boots                 instead of an OS          as an application
         every EC2 instance, every Azure VM and every GCE instance
         is a guest on a TYPE 1 hypervisor. The whole cloud is this
         experiment, at rack scale -- which is why it is experiment 1

      a 32 GB / 8 vCPU host, four guests:
      vm           RAM  vCPU   active
      web-1         8 G     4     35%
      web-2         8 G     4     30%
      db-1         16 G     4     90%
      batch        16 G     8     20%

      allocated RAM  : 48 GB on a 32 GB host (1.50x)
      allocated vCPU : 20 on 8 (2.50x)
      RAM actually touched     : 22.8 GB
      reclaimed by ballooning  : 25.2 GB
      swapping                 : 0.0 GB
         48 GB allocated on a 32 GB host and
         20 vCPUs on 8, and nothing is swapping -- because
         the guests only TOUCH 22.8 GB.
         Overcommit works on the same bet an airline makes, and it is
         why a cloud provider can sell more capacity than it owns

      now the batch job wakes up (20% -> 95% active):
      RAM actually touched : 34.8 GB
      swapping             : 2.8 GB
         2.8 GB OVER, AND NOW EVERY GUEST IS SLOW -- not just
         the batch job. Memory overcommit fails as a CLIFF, and it
         fails for the neighbours: this is the 'noisy neighbour'
         problem, and it is why cloud instance types quote DEDICATED
         memory and only burstable CPU.
         CPU overcommit degrades gracefully because time-slicing
         shares; RAM does not, because a page is either resident or
         it is not

    --- experiment 2: host a page on the server (this RUNS)
      document root : a new temporary folder, cloud13b_www_...
      serving       : http://127.0.0.1, on a port the system chose  (a REAL server)
      GET /         -> 200, text/html, 225 bytes
      GET /data.json-> 200, application/json, {'South': 10360.0, 'North': 2520.0}
      GET /missing  -> 404
         a page was written to a document root, served over TCP,
         fetched back, and its CONTENT-TYPE checked. That is the whole
         of experiment 2; Apache under XAMPP adds virtual hosts,
         .htaccess, PHP and TLS, and the shape is identical.
         Note the Content-Type header. A browser renders index.html
         because the server SAID text/html -- get that wrong and the
         browser downloads your page instead of showing it, which is
         the commonest 'my site is broken' on a fresh VM

      concern                 on your VM                managed (S3/App Service)
      who patches Apache      YOU, monthly              the provider
      TLS certificate         certbot, renewals         issued and rotated
      scaling                 a bigger VM               automatic
      a static site costs     a VM, hourly              cents per GB stored
      you control             everything                very little
         a STATIC site on a VM is the clearest case of paying for
         a general-purpose computer to do something an object store
         does for cents. Hosting index.html on S3 + CloudFront costs
         less than the VM's first hour

    --- experiment 7: the notebook environment
      In  [1]: import pandas as pd; import fixtures as f
      Out [1]: ok
      In  [2]: df.shape
      Out [2]: (9, 19)
      In  [3]: df.groupby('region')['revenue'].sum().to_dict()
      Out [3]: {'North': 2520.0, 'South': 10360.0}
      In  [4]: df['revenue'].sum()
      Out [4]: 12880.0
         four cells, executed in order, every output asserted.
         That is what a notebook TEST looks like -- papermill or
         nbconvert --execute do exactly this in CI, and a notebook
         nobody executes in CI is a notebook that has already
         drifted

                              Colab                   notebook on a cloud VM
      costs                   free tier, then paid    the INSTANCE, hourly
      data access             upload, or mount Drive  IAM role, no keys
      stops when              idle ~90 min            NEVER -- you stop it
      state on stop           LOST                    kept on the EBS volume
      GPU                     when available          the one you pay for
      private data            a policy question       inside your VPC

      an m5.xlarge notebook left running: $140.16/month
         'STOPS WHEN: NEVER' is the row that costs money. Colab
         disconnecting is an annoyance; a cloud notebook not
         disconnecting is a bill. Set an idle-shutdown lifecycle
         policy on day one -- SageMaker supports one, and it is the
         single most useful thing you can configure

Overcommit, measured. A 32 GB / 8 vCPU host with four guests:

allocated RAM  : 48 GB on a 32 GB host   (1.50x)
allocated vCPU : 20 on 8                 (2.50x)
RAM actually touched     : 22.8 GB
reclaimed by ballooning  : 25.2 GB
swapping                 : 0.0 GB

48 GB allocated on 32 GB, and nothing is swapping, because the guests only touch 22.8 GB. Overcommit works on the same bet an airline makes.

Then the batch job wakes up (20% → 95% active):

RAM actually touched : 34.8 GB
swapping             : 2.8 GB

Now every guest is slow — not just the batch job.

THE ASYMMETRY TO REMEMBER

NOTE

CPU overcommit degrades gracefully. Memory overcommit fails as a cliff.

CPU is time-sliced, so twice the demand means half the speed for everyone. A memory page is either resident or on disk, and the difference is a factor of thousands. That is the "noisy neighbour" problem, and it is why cloud instance types quote dedicated memory and only burstable CPU.

Changed: the program printed the temporary document root's full name and the web server's port, which the system picks afresh on every run; it now says what they are.

RESULT

48 GB allocated on a 32 GB host swaps nothing while the guests touch 22.8 GB; when the batch job wakes, 2.8 GB swaps and every guest slows.

Experiment 2 — Host a page on the server

1. Question

Install and configure Apache (or XAMPP) on the VM, and host a page.

2. Aim

Serve a page, check the header that decides whether a browser shows it, and compare a VM with an object store.

3. Steps

The procedure, on the console and CLI, 02_web_server.md:

  1. Install Apache, on Linux.
  2. Or install XAMPP, on Windows.
  3. Configure it.
  4. Enable TLS.
  5. Compare it with an object store.

WHAT RUNS

The page is really served, in 01_vm_and_hosting.py's experiment 2:

GET /          -> 200, text/html, 225 bytes
GET /data.json -> 200, application/json, {'South': 10360.0, 'North': 2520.0}
GET /missing   -> 404

A page was written to a document root, served over TCP, fetched back, and its Content-Type checked. That is the whole of experiment 2; Apache under XAMPP adds virtual hosts, .htaccess, PHP and TLS, and the shape is identical.

4. Programme

The procedure, on the console and CLI, 02_web_server.md:

# Experiment 2 -- install and configure Apache/XAMPP on the VM and host a page

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`01_vm_and_hosting.py`, which really does serve a page over TCP and fetch it back**.

---

<!-- Step 1: Install Apache, on Linux -->
## Linux (Apache directly)

```bash
sudo apt update && sudo apt install -y apache2
sudo systemctl enable --now apache2
systemctl status apache2

sudo tee /var/www/html/index.html >/dev/null <<'EOF'
<!doctype html>
<title>Sales dashboard</title>
<h1>Retail sales</h1>
EOF

curl -I http://localhost/          # 200, and check the Content-Type
```

From the HOST machine, using the guest's IP from experiment 1:

```
http://192.168.x.x/
```

**If that times out**, work through in this order:

1. `sudo ufw allow 80/tcp` — the guest firewall
2. the network mode: **NAT means the LAN cannot reach the guest** (see
   experiment 1). Switch to Bridged, or forward a port.
3. `sudo ss -tlnp | grep :80` — is Apache bound to `0.0.0.0` or only to
   `127.0.0.1`?

<!-- Step 2: Or install XAMPP, on Windows -->
## Windows (XAMPP/WAMP)

Install XAMPP, start **Apache** from the control panel, put the page in
`C:\xampp\htdocs\`, browse to `http://localhost/`.

**Apache will not start** almost always because **port 80 is taken** — by IIS,
by Skype (historically), or by another web server. Change
`Listen 80` to `Listen 8080` in `httpd.conf` and browse to
`http://localhost:8080/`.

<!-- Step 3: Configure it -->
## The configuration worth understanding

| Directive | Where | What it does |
|---|---|---|
| `DocumentRoot` | `apache2.conf` / `httpd.conf` | which directory is served |
| `Listen` | `ports.conf` | which port |
| `<VirtualHost>` | `sites-available/` | several sites on one IP, by hostname |
| `DirectoryIndex` | `apache2.conf` | which file `/` serves |
| `AddType` | `mime.types` | the **Content-Type** header |

**`AddType` is the one that bites.** A browser renders your page because the
server *said* `text/html`. Serve it as `text/plain` and the browser shows
source; serve it as `application/octet-stream` and the browser downloads it.
The runnable half asserts this header for exactly that reason.

<!-- Step 4: Enable TLS -->
## Enabling TLS

```bash
sudo apt install -y certbot python3-certbot-apache
sudo certbot --apache -d example.com     # needs a real domain pointing here
```

**You cannot get a certificate for `192.168.1.50`.** Certificate authorities
issue for names they can validate, so a VM on your LAN gets a self-signed
certificate and a browser warning — which is the correct behaviour, not a bug.

<!-- Step 5: Compare it with an object store -->
## The comparison that ends the experiment

| | This VM | S3 + CloudFront |
|---|---|---|
| Patch Apache | **you, monthly** | not your problem |
| TLS certificate | certbot, and renewals | issued and rotated |
| Survives your laptop closing | **no** | yes |
| Handles a traffic spike | no | yes |
| Cost for a static page | a VM, hourly | **cents per GB** |

**A static site on a VM is a general-purpose computer doing an object store's
job.** That is the point the experiment makes by having you do it the hard way
once.

5. Execution and Results

The procedure, on the console and CLI, 02_web_server.md:

NOT RUN HERE

02_web_server.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

THE HEADER THAT DECIDES WHETHER YOUR SITE WORKS

A browser renders index.html because the server said text/html. Get AddType wrong and the browser downloads your page instead of showing it — the commonest "my site is broken" on a fresh VM, and the reason the runnable half asserts the header rather than just the body.

And the comparison that ends the experiment:

This VM S3 + CloudFront
Patch Apache you, monthly not your problem
TLS certificate certbot, and renewals issued and rotated
Survives your laptop closing no yes
A static page costs a VM, hourly cents per GB

A static site on a VM is a general-purpose computer doing an object store's job, and the experiment makes the point by having you do it the hard way once.

The Python model for this experiment, 01_vm_and_hosting.py, is shown in full under Experiment 1, with what it printed.

RESULT

A page written to a document root was served over TCP and fetched back as text/html; a missing page gave 404.

Experiment 3 — Create and configure a cloud account

1. Question

Create and configure a cloud account on the AWS, Azure or GCP free tier.

2. Aim

Set the account up safely, and see exactly how IAM decides each request.

3. Steps

The procedure, on the console and CLI, 03_account_setup.md:

  1. Do the six things first.
  2. Use an IAM user, not root.
  3. Stay within the free tier.
  4. Know the equivalents.
  5. Clean up.

The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:

  1. Evaluate the attached policies.
  2. Add full S3 admin.
  3. Reverse the policy order.
  4. Compare a role with a user.
  5. Compare least privilege with :.
  6. Price the free tier.

THE THREE RULES, AND THEY ARE THE WHOLE SUBJECT

  1. An explicit DENY anywhere wins — always, unconditionally.
  2. Otherwise, an ALLOW that matches grants access.
  3. Otherwise DENY — the implicit deny.

Evaluated against a realistic policy set:

Action Resource Result Why
s3:GetObject retail-lake/raw/sales.csv Allow DataScientistRead
s3:PutObject retail-lake/raw/sales.csv Deny explicit deny in ProtectRawZone
s3:PutObject retail-lake/models/model.pkl Allow SageMakerExecution
s3:GetObject other-bucket/secret.csv Deny implicit — nothing matched

Read rows 2 and 3 together. The same action on the same bucket is denied under raw/ and allowed under models/, because a Deny scoped to one prefix beats an Allow scoped to the bucket. That is how a data lake keeps a raw zone immutable while the rest stays writable.

4. Programme

The procedure, on the console and CLI, 03_account_setup.md:

# Experiment 3 -- create and configure a cloud account (AWS/Azure/GCP free tier)

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`03_iam_and_account.py`, which implements and exercises IAM's evaluation algorithm**.

---

<!-- Step 1: Do the six things first -->
## Do these six things in this order, before anything else

1. **Enable MFA on the root account.** Then never use the root account again.
2. **Create an IAM user for yourself** with `AdministratorAccess`, and use
   that.
3. **Set a budget alarm.** Billing → Budgets → a monthly cost budget with an
   alert at 50%, 80% and 100%. **Do this on day one**, not after the bill.
4. **Enable Cost Explorer.** It takes 24 hours to populate, so turn it on now.
5. **Pick one region and stay in it.** Resources in another region are
   invisible in the console and still bill.
6. **Install and configure the CLI:**

```bash
aws configure          # access key, secret, default region, output format
aws sts get-caller-identity        # who am I, really
aws configure list-profiles
```

<!-- Step 2: Use an IAM user, not root -->
## Root against IAM user

| | Root | IAM user |
|---|---|---|
| Created by | signing up | you |
| Can | **everything, including closing the account** | what its policies allow |
| MFA | **mandatory, immediately** | strongly recommended |
| Daily use | **never** | yes |
| Recovers | account-level problems only | — |

<!-- Step 3: Stay within the free tier -->
## The free tier, honestly

| Type | Duration | Examples |
|---|---|---|
| **12 months free** | from signup | 750 h/month t3.micro, 5 GB S3, 750 h RDS |
| **Always free** | forever | 1 M Lambda requests/month, 25 GB DynamoDB |
| **Trials** | short, one-off | SageMaker 250 h for 2 months |

**Three things bill you anyway, inside the free tier:**

- **Egress.** Ingress is free; downloading is not.
- **Anything larger than the free size.** A `t3.small` is not a `t3.micro`.
- **Resources you forgot.** An idle NAT Gateway is about $32/month, an
  unattached Elastic IP is about $3.60/month, and a SageMaker endpoint is
  about $70/month.

<!-- Step 4: Know the equivalents -->
## The equivalents, when the exam asks

| Concept | AWS | Azure | GCP |
|---|---|---|---|
| Object storage | S3 | Blob Storage | Cloud Storage |
| Block storage | EBS | Managed Disks | Persistent Disk |
| File storage | EFS | Azure Files | Filestore |
| VM | EC2 | Virtual Machines | Compute Engine |
| Managed SQL | RDS | SQL Database | Cloud SQL |
| Warehouse | Redshift | Synapse | **BigQuery** |
| Serverless functions | Lambda | Functions | Cloud Functions |
| Managed ML | SageMaker | Azure ML | Vertex AI |
| Monitoring | CloudWatch | Monitor | Cloud Monitoring |
| Identity | IAM | Entra ID | IAM |

**Learn the row, not the column.** Every provider has all of these, and an
exam answer that names the concept and gives one vendor's term for it is
worth more than a memorised AWS product list.

<!-- Step 5: Clean up -->
## Then clean up

```bash
aws ec2 describe-instances --query \
  'Reservations[].Instances[?State.Name==`running`].[InstanceId,InstanceType]'
aws s3 ls
aws sagemaker list-endpoints
```

**Run those three at the end of every lab session.** Nothing else in this
course will save you as much money.

The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:

"""Experiments 3 and 10 -- cloud account setup, and IAM roles for SageMaker.

There is no cloud account for this repository and none will be created, so
`03_account_setup.md` and `10_sagemaker_notebook.md` carry the console
click-paths, marked as not run here.

What runs here is the part that is actually examinable: AWS's policy
evaluation algorithm, implemented in iam.py and exercised against a realistic
policy set. The three rules are the whole subject.
"""
import fixtures as f
from iam import evaluate


def show(policies, cases, title):
    print(f"\n    {title}")
    print(f"      {'action':<32}{'resource':<47}{'result':<7}why")
    for action, resource in cases:
        trace = {}
        decision = evaluate(policies, action, resource, trace)
        print(f"      {action:<32}{resource:<47}"
              f"{decision:<7}{trace['reason']}")
    return {(a, r): evaluate(policies, a, r) for a, r in cases}


def main():
    print("  Experiments 3 and 10 -- accounts, roles and IAM evaluation")

    print("""
    the three rules, in order:
      1. an EXPLICIT DENY anywhere wins -- always, unconditionally
      2. otherwise an ALLOW that matches grants access
      3. otherwise DENY -- the IMPLICIT DENY""")

    # Step 1: Evaluate the attached policies
    cases = [
        ("s3:GetObject",    "arn:aws:s3:::retail-lake/raw/sales.csv"),
        ("s3:PutObject",    "arn:aws:s3:::retail-lake/raw/sales.csv"),
        ("s3:PutObject",    "arn:aws:s3:::retail-lake/models/model.pkl"),
        ("s3:DeleteObject", "arn:aws:s3:::retail-lake/raw/sales.csv"),
        ("s3:GetObject",    "arn:aws:s3:::other-bucket/secret.csv"),
        ("sagemaker:CreateTrainingJob", "arn:aws:sagemaker:*:*:training-job/x"),
    ]
    got = show(f.POLICIES, cases, "the attached policies, evaluated:")

    assert got[("s3:GetObject", "arn:aws:s3:::retail-lake/raw/sales.csv")] == "Allow"
    assert got[("s3:PutObject", "arn:aws:s3:::retail-lake/raw/sales.csv")] == "Deny"
    assert got[("s3:PutObject", "arn:aws:s3:::retail-lake/models/model.pkl")] == "Allow"
    assert got[("s3:GetObject", "arn:aws:s3:::other-bucket/secret.csv")] == "Deny"

    print("""         READ ROWS 2 AND 3 TOGETHER. The same action on the same
         bucket is denied under raw/ and allowed under models/, because
         a Deny statement scoped to one prefix beats an Allow scoped to
         the bucket. Prefix-scoped policies are how a data lake keeps a
         raw zone immutable while the rest stays writable""")

    # Step 2: Add full S3 admin
    print("\n    now ADD a policy granting s3:* on everything:")
    admin = {"name": "S3FullAccess",
             "statements": [{"Effect": "Allow", "Action": ["s3:*"],
                             "Resource": ["*"]}]}
    with_admin = f.POLICIES + [admin]
    trace = {}
    still = evaluate(with_admin, "s3:PutObject",
                     "arn:aws:s3:::retail-lake/raw/sales.csv", trace)
    print(f"      s3:PutObject on raw/  -> {still}   ({trace['reason']})")
    assert still == "Deny"
    print("""         STILL DENIED, with full S3 admin attached. An explicit
         Deny cannot be out-voted, out-numbered or out-scoped -- there
         is no 'more specific allow wins' rule. To lift it you must
         REMOVE the Deny.
         This is the single most common IAM misunderstanding, and it is
         also the feature: a Deny is how an organisation guarantees
         something, rather than hoping nobody granted otherwise""")

    # Step 3: Reverse the policy order
    reversed_order = list(reversed(with_admin))
    assert evaluate(reversed_order, "s3:PutObject",
                    "arn:aws:s3:::retail-lake/raw/sales.csv") == "Deny"
    print("\n    policy ORDER does not matter -- reversed, the answer is the same")
    print("""         unlike a firewall rule list, IAM is not first-match.
         Every statement is evaluated, then the three rules decide.
         Say that and you have answered 'how does IAM resolve
         conflicting policies?'""")

    # Step 4: Compare a role with a user
    print("\n    a ROLE is not a user:")
    print(f"      {'':<16}{'user':<30}{'role'}")
    for label, u, r in (
            ("credentials", "long-lived access key", "TEMPORARY, auto-rotated"),
            ("who assumes it", "a person", "a SERVICE or another principal"),
            ("in a notebook", "keys in a file  <- BAD", "attached; no keys exist"),
            ("if leaked", "valid until revoked", "expires in minutes to hours")):
        print(f"      {label:<16}{u:<30}{r}")
    print("""         a SageMaker notebook gets an EXECUTION ROLE, so no access
         key is ever written to disk. That is why experiment 10 says
         'attach IAM role' rather than 'paste your credentials', and
         'I put my keys in the notebook' is the answer that loses the
         marks""")

    # Step 5: Compare least privilege with *:*
    print("\n    least privilege, as an exercise:")
    over = {"name": "ItWorksNow",
            "statements": [{"Effect": "Allow", "Action": ["*"],
                            "Resource": ["*"]}]}
    tight = {"name": "TrainingJobOnly",
             "statements": [
                 {"Effect": "Allow",
                  "Action": ["s3:GetObject"],
                  "Resource": ["arn:aws:s3:::retail-lake/train/*"]},
                 {"Effect": "Allow",
                  "Action": ["s3:PutObject"],
                  "Resource": ["arn:aws:s3:::retail-lake/models/*"]},
             ]}
    probe = [("s3:GetObject", "arn:aws:s3:::retail-lake/train/x.csv"),
             ("s3:PutObject", "arn:aws:s3:::retail-lake/models/m.tar.gz"),
             ("iam:CreateUser", "*"),
             ("ec2:TerminateInstances", "*")]
    print(f"      {'action':<26}{'*:* policy':<14}{'scoped policy'}")
    for a, r in probe:
        print(f"      {a:<26}{evaluate([over], a, r):<14}"
              f"{evaluate([tight], a, r)}")
    assert evaluate([over], "iam:CreateUser", "*") == "Allow"
    assert evaluate([tight], "iam:CreateUser", "*") == "Deny"
    print("""         both policies let the training job run. One of them also
         lets it create IAM users and terminate every instance in the
         account. '*:* made it work' is not a solution, it is a
         postponed incident""")

    # Step 6: Price the free tier
    print("\n    the free tier, and the three things that bill anyway:")
    print(f"      {'service':<22}{'free tier':<34}{'what still costs'}")
    for svc, free, cost in (
            ("EC2", "750 hrs/month t2/t3.micro, 12 mo", "any larger instance"),
            ("S3", "5 GB Standard, 12 mo", "EGRESS to the internet"),
            ("RDS", "750 hrs/month db.t3.micro, 12 mo", "storage over 20 GB"),
            ("Lambda", "1M requests/month, ALWAYS free", "duration x memory"),
            ("SageMaker", "250 hrs notebook, 2 mo", "ENDPOINTS, billed hourly")):
        print(f"      {svc:<22}{free:<34}{cost}")
    endpoint_month = f.EC2["m5.large"] * f.HOURS_PER_MONTH
    print(f"\n      a forgotten ml.m5.large endpoint costs about "
          f"${endpoint_month:,.0f}/month")
    assert 60 < endpoint_month < 90
    print("""         THE ENDPOINT IS THE TRAP. A training job ends and stops
         billing; an endpoint runs until you delete it, at hourly
         rates, whether or not anything calls it. Every 'I got a
         surprise AWS bill' story is a resource nobody switched off --
         set a BUDGET ALARM on day one, before anything else""")


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 03_account_setup.md:

NOT RUN HERE

03_account_setup.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 03_iam_and_account.py, for experiments 3 and 10:

OUTPUT

  Experiments 3 and 10 -- accounts, roles and IAM evaluation

    the three rules, in order:
      1. an EXPLICIT DENY anywhere wins -- always, unconditionally
      2. otherwise an ALLOW that matches grants access
      3. otherwise DENY -- the IMPLICIT DENY

    the attached policies, evaluated:
      action                          resource                                       result why
      s3:GetObject                    arn:aws:s3:::retail-lake/raw/sales.csv         Allow  Allow in DataScientistRead
      s3:PutObject                    arn:aws:s3:::retail-lake/raw/sales.csv         Deny   EXPLICIT DENY in ProtectRawZone
      s3:PutObject                    arn:aws:s3:::retail-lake/models/model.pkl      Allow  Allow in SageMakerExecution
      s3:DeleteObject                 arn:aws:s3:::retail-lake/raw/sales.csv         Deny   EXPLICIT DENY in ProtectRawZone
      s3:GetObject                    arn:aws:s3:::other-bucket/secret.csv           Deny   IMPLICIT DENY -- no statement matched
      sagemaker:CreateTrainingJob     arn:aws:sagemaker:*:*:training-job/x           Allow  Allow in SageMakerExecution
         READ ROWS 2 AND 3 TOGETHER. The same action on the same
         bucket is denied under raw/ and allowed under models/, because
         a Deny statement scoped to one prefix beats an Allow scoped to
         the bucket. Prefix-scoped policies are how a data lake keeps a
         raw zone immutable while the rest stays writable

    now ADD a policy granting s3:* on everything:
      s3:PutObject on raw/  -> Deny   (EXPLICIT DENY in ProtectRawZone)
         STILL DENIED, with full S3 admin attached. An explicit
         Deny cannot be out-voted, out-numbered or out-scoped -- there
         is no 'more specific allow wins' rule. To lift it you must
         REMOVE the Deny.
         This is the single most common IAM misunderstanding, and it is
         also the feature: a Deny is how an organisation guarantees
         something, rather than hoping nobody granted otherwise

    policy ORDER does not matter -- reversed, the answer is the same
         unlike a firewall rule list, IAM is not first-match.
         Every statement is evaluated, then the three rules decide.
         Say that and you have answered 'how does IAM resolve
         conflicting policies?'

    a ROLE is not a user:
                      user                          role
      credentials     long-lived access key         TEMPORARY, auto-rotated
      who assumes it  a person                      a SERVICE or another principal
      in a notebook   keys in a file  <- BAD        attached; no keys exist
      if leaked       valid until revoked           expires in minutes to hours
         a SageMaker notebook gets an EXECUTION ROLE, so no access
         key is ever written to disk. That is why experiment 10 says
         'attach IAM role' rather than 'paste your credentials', and
         'I put my keys in the notebook' is the answer that loses the
         marks

    least privilege, as an exercise:
      action                    *:* policy    scoped policy
      s3:GetObject              Allow         Allow
      s3:PutObject              Allow         Allow
      iam:CreateUser            Allow         Deny
      ec2:TerminateInstances    Allow         Deny
         both policies let the training job run. One of them also
         lets it create IAM users and terminate every instance in the
         account. '*:* made it work' is not a solution, it is a
         postponed incident

    the free tier, and the three things that bill anyway:
      service               free tier                         what still costs
      EC2                   750 hrs/month t2/t3.micro, 12 mo  any larger instance
      S3                    5 GB Standard, 12 mo              EGRESS to the internet
      RDS                   750 hrs/month db.t3.micro, 12 mo  storage over 20 GB
      Lambda                1M requests/month, ALWAYS free    duration x memory
      SageMaker             250 hrs notebook, 2 mo            ENDPOINTS, billed hourly

      a forgotten ml.m5.large endpoint costs about $70/month
         THE ENDPOINT IS THE TRAP. A training job ends and stops
         billing; an endpoint runs until you delete it, at hourly
         rates, whether or not anything calls it. Every 'I got a
         surprise AWS bill' story is a resource nobody switched off --
         set a BUDGET ALARM on day one, before anything else

NOW ADD FULL S3 ADMIN

add a policy granting s3:* on *
s3:PutObject on raw/  ->  Deny   (EXPLICIT DENY in ProtectRawZone)

Still denied. An explicit Deny cannot be out-voted, out-numbered or out-scoped — there is no "more specific allow wins" rule. To lift it you must remove the Deny.

This is the single most common IAM misunderstanding, and it is also the feature: a Deny is how an organisation guarantees something rather than hoping nobody granted otherwise.

And policy order does not matter — reversed, the answer is identical. Unlike a firewall rule list, IAM is not first-match: every statement is evaluated, then the three rules decide.

Least privilege, made concrete:

Action *:* policy scoped policy
s3:GetObject on train/ Allow Allow
s3:PutObject on models/ Allow Allow
iam:CreateUser Allow Deny
ec2:TerminateInstances Allow Deny

Both policies let the training job run. One of them also lets it create IAM users and terminate every instance in the account. "*:* made it work" is not a solution, it is a postponed incident.

RESULT

An explicit Deny on raw/ beats every Allow, even full S3 admin; policy order changes nothing; a scoped policy runs the job and refuses everything else.

Experiment 4 — Create and manage storage buckets

1. Question

Create and manage storage buckets, and upload and access datasets.

2. Aim

Store and list objects, and see that an object store has no directories, no rename and a bill for every version.

3. Steps

The procedure, on the console and CLI, 04_buckets.md:

  1. Create a bucket and copy data, with the CLI.
  2. Name it uniquely.
  3. Turn on versioning and a lifecycle rule.
  4. Block public access.
  5. Encrypt it.

The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:

  1. Experiment 4: list a bucket by prefix.
  2. Rename an object.
  3. Version it.
  4. Price the storage classes.
  5. Read it twice a month.
  6. Price the egress.
  7. Experiments 5 and 6: compare block, file and object storage.
  8. Provision against consumption.

THERE ARE NO DIRECTORIES

LIST prefix 'raw/' with delimiter '/':
  objects at this level : (none)
  common prefixes       : ['raw/2026/']

raw/2026/01/sales.csv is ONE KEY containing three slashes. The console's folder tree is drawn from common prefixes computed at list time. Delete every object under a "folder" and the folder is gone, because it never existed.

And a prefix scan is the only query an object store supports.

4. Programme

The procedure, on the console and CLI, 04_buckets.md:

# Experiment 4 -- create and manage storage buckets; upload and access datasets

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`04_storage.py`, which implements the key semantics and runs the pricing arithmetic**.

---

<!-- Step 1: Create a bucket and copy data, with the CLI -->
## The CLI, which is faster than the console

```bash
aws s3 mb s3://retail-lake-<your-unique-suffix>     # names are GLOBAL
aws s3 cp sales.csv s3://retail-lake/raw/2026/01/
aws s3 cp ./data s3://retail-lake/raw/ --recursive
aws s3 ls s3://retail-lake/raw/                    # by PREFIX
aws s3 ls s3://retail-lake/raw/ --recursive --human-readable --summarize
aws s3 sync ./local s3://retail-lake/curated/      # only what changed
aws s3 mv s3://retail-lake/a.csv s3://retail-lake/b.csv   # COPY + DELETE
aws s3 rm s3://retail-lake/raw/old.csv
aws s3 presign s3://retail-lake/raw/sales.csv --expires-in 3600
```

**`s3 mv` is not a rename.** There is no rename. It copies the whole object
and deletes the original, and you are billed for both.

<!-- Step 2: Name it uniquely -->
## Bucket names are globally unique

Across every AWS customer on earth. `s3://data` was taken in 2006. Use
`<org>-<purpose>-<region>-<random>`.

<!-- Step 3: Turn on versioning and a lifecycle rule -->
## Versioning and lifecycle

```bash
aws s3api put-bucket-versioning --bucket retail-lake \
    --versioning-configuration Status=Enabled
```

**Then set a lifecycle rule immediately**, or you pay for every version of
every object for ever:

```json
{"Rules": [{
  "ID": "tier-and-expire",
  "Filter": {"Prefix": "raw/"},
  "Status": "Enabled",
  "Transitions": [
    {"Days": 30,  "StorageClass": "STANDARD_IA"},
    {"Days": 90,  "StorageClass": "GLACIER"}
  ],
  "NoncurrentVersionExpiration": {"NoncurrentDays": 30},
  "AbortIncompleteMultipartUpload": {"DaysAfterInitiation": 7}
}]}
```

**That last rule matters more than it looks.** A failed multipart upload
leaves parts that are invisible in `s3 ls` and billed for ever. Many
mysterious S3 bills are abandoned multipart uploads.

<!-- Step 4: Block public access -->
## Blocking public access

```bash
aws s3api put-public-access-block --bucket retail-lake \
  --public-access-block-configuration \
  "BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true"
```

**On by default since 2023**, and it should stay on. "Public S3 bucket" was
the single most common cause of data breaches for a decade. To share one
object, use a **presigned URL** — time-limited, no policy change.

<!-- Step 5: Encrypt it -->
## Encryption

Server-side encryption is on by default (SSE-S3). Use **SSE-KMS** when you
need an audit trail of who decrypted what, and accept that KMS charges per
request. `--sse aws:kms --sse-kms-key-id <arn>`.

The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:

"""Experiments 4, 5 and 6 -- object, block and file storage on the cloud.

The console click-paths for S3, EBS and EFS are in `04_buckets.md`,
`05_ebs.md` and `06_efs.md`, all marked as not run here -- there is no cloud
account.

What runs here is the SEMANTICS, which is what the exam is about: why an
object store has no directories, why you cannot rename, why a block volume
attaches to exactly one instance, and the pricing arithmetic that decides
which one you should have chosen.
"""
import fixtures as f
from objectstore import ObjectStore


def money(x):
    return f"${x:,.2f}"


def main():
    print("  Experiments 4, 5 and 6 -- object, block and file storage")

    # Step 1: Experiment 4: list a bucket by prefix
    # ================================================== experiment 4
    print("\n    --- experiment 4: buckets and objects")
    s3 = ObjectStore("retail-lake")
    keys = [
        "raw/2026/01/sales.csv",
        "raw/2026/02/sales.csv",
        "raw/2026/03/sales.csv",
        "curated/2026/sales.parquet",
        "models/churn/model.tar.gz",
        "README.md",
    ]
    for k in keys:
        s3.put(k, b"x" * 1024)
    print(f"      {len(s3.objects)} objects, {s3.total_bytes():,} bytes")

    print("\n      LIST with prefix 'raw/' and delimiter '/':")
    plain, prefixes = s3.list("raw/", delimiter="/")
    print(f"        objects at this level : {plain or '(none)'}")
    print(f"        common prefixes       : {prefixes}")
    assert plain == [] and prefixes == ["raw/2026/"]
    print("""         THERE IS NO DIRECTORY ANYWHERE. 'raw/2026/01/sales.csv' is
         ONE KEY containing three slashes, and the console's folder
         tree is drawn from COMMON PREFIXES computed at list time.
         Delete every object under a 'folder' and the folder is gone,
         because it never existed""")

    print("\n      LIST with prefix 'raw/' and NO delimiter:")
    plain, prefixes = s3.list("raw/")
    for k in plain:
        print(f"        {k}")
    assert len(plain) == 3 and prefixes == []
    print("""         without a delimiter you get every key under the prefix,
         flat. That is the only query an object store supports: a
         PREFIX SCAN. No WHERE, no index, no search by content""")

    # Step 2: Rename an object
    cost = s3.rename("README.md", "docs/README.md")
    print(f"\n      'rename' README.md -> docs/README.md")
    print(f"        bytes read {cost['bytes_read']:,}, "
          f"bytes written {cost['bytes_written']:,}, "
          f"API calls {cost['requests']}")
    assert "README.md" not in s3.objects and "docs/README.md" in s3.objects
    print("""         THERE IS NO RENAME. It is a COPY plus a DELETE, so it
         reads and writes the whole object. Renaming a 5 TB dataset
         'to tidy up the folders' moves 10 TB and is billed for it --
         and on a filesystem it would have been a metadata edit""")

    # Step 3: Version it
    print("\n      versioning ON, then overwrite and delete:")
    s3.versioning = True
    s3.put("curated/2026/sales.parquet", b"y" * 2048)
    s3.delete("curated/2026/sales.parquet")
    hist = s3.versions["curated/2026/sales.parquet"]
    current = s3.objects["curated/2026/sales.parquet"]
    print(f"        older versions kept : {len(hist)}")
    print(f"        current object      : {current['class']}")
    assert len(hist) == 2 and current["class"] == "DeleteMarker"
    print("""         a DELETE with versioning on writes a DELETE MARKER; the
         data is still there and still billed. That is the feature
         (you can undelete) and the bill (you are paying for every
         version of every object until a lifecycle rule removes
         them)""")

    # Step 4: Price the storage classes
    print("\n    storage classes -- 1 TB stored for a year, "
          "retrieved once:")
    gb = 1024
    print(f"      {'class':<22}{'storage/yr':>12}{'retrieve 1 TB':>15}"
          f"{'total':>12}{'min days':>10}")
    rows = {}
    for cls in f.S3_STORAGE:
        store = f.S3_STORAGE[cls] * gb * 12
        retrieve = f.S3_RETRIEVAL[cls] * gb
        rows[cls] = store + retrieve
        print(f"      {cls:<22}{money(store):>12}{money(retrieve):>15}"
              f"{money(store + retrieve):>12}{f.S3_MIN_DAYS[cls]:>10}")
    ratio = rows["Standard"] / rows["Glacier Deep Archive"]
    store_only = f.S3_STORAGE["Standard"] / f.S3_STORAGE["Glacier Deep Archive"]
    assert rows["Glacier Deep Archive"] < rows["Standard"] / 5
    assert store_only > 2 * ratio
    print(f"""         READ THE TWO RATIOS SEPARATELY. Deep Archive STORAGE is
         {store_only:.0f}x cheaper than Standard -- but add ONE retrieval a year
         and the all-in saving falls to {ratio:.1f}x, because the retrieval
         fee ({money(f.S3_RETRIEVAL['Glacier Deep Archive'] * gb)}) is larger than a year of its storage
         ({money(f.S3_STORAGE['Glacier Deep Archive'] * gb * 12)}). The headline discount is not the discount.
         And it has a 180-DAY MINIMUM BILLING DURATION plus a
         retrieval that takes up to 12 hours: delete an object after
         10 days and you are billed for 180.
         Storage class is a bet on your ACCESS PATTERN, and the
         penalties are what make the bet real""")

    # Step 5: Read it twice a month
    print("\n    the same 1 TB, retrieved TWICE A MONTH:")
    print(f"      {'class':<22}{'storage/yr':>12}{'retrieval/yr':>14}{'total':>12}")
    freq = {}
    for cls in ("Standard", "Standard-IA", "Glacier Instant"):
        store = f.S3_STORAGE[cls] * gb * 12
        retrieve = f.S3_RETRIEVAL[cls] * gb * 24
        freq[cls] = store + retrieve
        print(f"      {cls:<22}{money(store):>12}{money(retrieve):>14}"
              f"{money(store + retrieve):>12}")
    assert freq["Standard"] < freq["Standard-IA"]
    assert freq["Standard"] < freq["Glacier Instant"]
    print("""         STANDARD IS NOW THE CHEAPEST. The 'cheap' tiers charge
         per GB retrieved, and at two retrievals a month the retrieval
         fee exceeds everything the storage discount saved.
         Infrequent-access tiers are for data you genuinely do not
         touch -- and 'we moved everything to IA to save money' is how
         a bill goes UP""")

    # Step 6: Price the egress
    print("\n    egress, which is the line item nobody predicts:")
    print(f"      {'transfer':<40}{'cost'}")
    for label, amount in (("1 TB in from the internet", 0),
                          ("1 TB out to the internet", gb * f.EGRESS_PER_GB),
                          ("1 TB between AZs", gb * 0.01 * 2),
                          ("1 TB S3 -> EC2, same region", 0)):
        print(f"      {label:<40}{money(amount)}")
    month_store = f.S3_STORAGE["Standard"] * gb
    one_egress = gb * f.EGRESS_PER_GB
    months = one_egress / month_store
    assert months > 3
    print(f"""         DOWNLOADING 1 TB ONCE ({money(one_egress)}) COSTS AS MUCH AS
         STORING IT FOR {months:.1f} MONTHS ({money(month_store)}/month). Ingress is
         free; egress is not. That asymmetry is the economic shape of
         every cloud -- cheap to put data in, expensive to take it out
         -- and it is the mechanism behind vendor lock-in: your data
         is not held hostage, it is simply expensive to move.
         It is also why 'move the compute to the data' survived from
         Course 12 B into the cloud era: run the job in the region
         holding the bucket and the transfer is free""")

    # Step 7: Experiments 5 and 6: compare block, file and object storage
    # ================================================== experiments 5, 6
    print("\n    --- experiments 5 and 6: block and file storage")
    print(f"\n      {'':<18}{'BLOCK (EBS)':<26}{'FILE (EFS)':<26}"
          f"{'OBJECT (S3)'}")
    for label, blk, fil, obj in (
            ("looks like", "a raw disk", "a mounted filesystem", "an HTTP API"),
            ("attached to", "ONE instance*", "MANY instances", "anything, anywhere"),
            ("access unit", "a 512 B block", "a file, byte ranges", "a whole object"),
            ("in-place edit", "yes", "yes", "NO -- rewrite it"),
            ("directories", "whatever the FS does", "real", "NONE -- prefixes"),
            ("latency", "sub-millisecond", "low millisecond", "tens of ms"),
            ("capacity", "provisioned, fixed", "elastic", "unlimited"),
            ("survives instance", "yes, if not root", "yes", "yes"),
            ("$/GB-month", f"{f.EBS_GP3_GB_MONTH:.3f}",
             f"{f.EFS_STANDARD_GB_MONTH:.2f}",
             f"{f.S3_STORAGE['Standard']:.3f}"),
    ):
        print(f"      {label:<18}{blk:<26}{fil:<26}{obj}")
    print("      * EBS Multi-Attach exists for io1/io2 and needs a "
          "cluster filesystem")

    print("\n    1 TB for a month, by type:")
    for label, price in (("EBS gp3", f.EBS_GP3_GB_MONTH),
                         ("EFS Standard", f.EFS_STANDARD_GB_MONTH),
                         ("S3 Standard", f.S3_STORAGE["Standard"])):
        print(f"      {label:<16}{money(price * gb):>10}")
    ebs, efs, s3c = (p * gb for p in (f.EBS_GP3_GB_MONTH,
                                      f.EFS_STANDARD_GB_MONTH,
                                      f.S3_STORAGE["Standard"]))
    assert efs > ebs > s3c
    print(f"""         EFS costs {efs / s3c:.0f}x what S3 does and {efs / ebs:.1f}x what EBS does,
         and it is worth it precisely when several instances must
         share a POSIX filesystem. Paying {efs / s3c:.0f}x for a dataset that one
         batch job reads once is the mistake -- that dataset belongs
         in S3.
         Choose by ACCESS PATTERN, not by price: block for a database
         or a boot disk, file for shared POSIX, object for everything
         a data pipeline reads""")

    # Step 8: Provision against consumption
    print("\n    provisioned against consumed, on a 1 TB EBS volume "
          "holding 200 GB:")
    print(f"      EBS billed on PROVISIONED size : "
          f"{money(f.EBS_GP3_GB_MONTH * gb)}/month")
    print(f"      S3 billed on STORED bytes      : "
          f"{money(f.S3_STORAGE['Standard'] * 200)}/month")
    over = f.EBS_GP3_GB_MONTH * gb
    used = f.S3_STORAGE["Standard"] * 200
    assert over > used * 15
    print(f"""         a factor of {over / used:.0f}. EBS bills the volume you asked for,
         empty or not; S3 bills the bytes you actually stored. An
         over-provisioned volume is invisible waste, which is why
         'just make it 1 TB to be safe' is an expensive habit""")


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 04_buckets.md:

NOT RUN HERE

04_buckets.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 04_storage.py, for experiments 4, 5 and 6:

OUTPUT

  Experiments 4, 5 and 6 -- object, block and file storage

    --- experiment 4: buckets and objects
      6 objects, 6,144 bytes

      LIST with prefix 'raw/' and delimiter '/':
        objects at this level : (none)
        common prefixes       : ['raw/2026/']
         THERE IS NO DIRECTORY ANYWHERE. 'raw/2026/01/sales.csv' is
         ONE KEY containing three slashes, and the console's folder
         tree is drawn from COMMON PREFIXES computed at list time.
         Delete every object under a 'folder' and the folder is gone,
         because it never existed

      LIST with prefix 'raw/' and NO delimiter:
        raw/2026/01/sales.csv
        raw/2026/02/sales.csv
        raw/2026/03/sales.csv
         without a delimiter you get every key under the prefix,
         flat. That is the only query an object store supports: a
         PREFIX SCAN. No WHERE, no index, no search by content

      'rename' README.md -> docs/README.md
        bytes read 1,024, bytes written 1,024, API calls 2
         THERE IS NO RENAME. It is a COPY plus a DELETE, so it
         reads and writes the whole object. Renaming a 5 TB dataset
         'to tidy up the folders' moves 10 TB and is billed for it --
         and on a filesystem it would have been a metadata edit

      versioning ON, then overwrite and delete:
        older versions kept : 2
        current object      : DeleteMarker
         a DELETE with versioning on writes a DELETE MARKER; the
         data is still there and still billed. That is the feature
         (you can undelete) and the bill (you are paying for every
         version of every object until a lifecycle rule removes
         them)

    storage classes -- 1 TB stored for a year, retrieved once:
      class                   storage/yr  retrieve 1 TB       total  min days
      Standard                   $282.62          $0.00     $282.62         0
      Standard-IA                $153.60         $10.24     $163.84        30
      One Zone-IA                $122.88         $10.24     $133.12        30
      Glacier Instant             $49.15         $30.72      $79.87        90
      Glacier Flexible            $44.24         $10.24      $54.48        90
      Glacier Deep Archive        $12.17         $20.48      $32.65       180
         READ THE TWO RATIOS SEPARATELY. Deep Archive STORAGE is
         23x cheaper than Standard -- but add ONE retrieval a year
         and the all-in saving falls to 8.7x, because the retrieval
         fee ($20.48) is larger than a year of its storage
         ($12.17). The headline discount is not the discount.
         And it has a 180-DAY MINIMUM BILLING DURATION plus a
         retrieval that takes up to 12 hours: delete an object after
         10 days and you are billed for 180.
         Storage class is a bet on your ACCESS PATTERN, and the
         penalties are what make the bet real

    the same 1 TB, retrieved TWICE A MONTH:
      class                   storage/yr  retrieval/yr       total
      Standard                   $282.62         $0.00     $282.62
      Standard-IA                $153.60       $245.76     $399.36
      Glacier Instant             $49.15       $737.28     $786.43
         STANDARD IS NOW THE CHEAPEST. The 'cheap' tiers charge
         per GB retrieved, and at two retrievals a month the retrieval
         fee exceeds everything the storage discount saved.
         Infrequent-access tiers are for data you genuinely do not
         touch -- and 'we moved everything to IA to save money' is how
         a bill goes UP

    egress, which is the line item nobody predicts:
      transfer                                cost
      1 TB in from the internet               $0.00
      1 TB out to the internet                $92.16
      1 TB between AZs                        $20.48
      1 TB S3 -> EC2, same region             $0.00
         DOWNLOADING 1 TB ONCE ($92.16) COSTS AS MUCH AS
         STORING IT FOR 3.9 MONTHS ($23.55/month). Ingress is
         free; egress is not. That asymmetry is the economic shape of
         every cloud -- cheap to put data in, expensive to take it out
         -- and it is the mechanism behind vendor lock-in: your data
         is not held hostage, it is simply expensive to move.
         It is also why 'move the compute to the data' survived from
         Course 12 B into the cloud era: run the job in the region
         holding the bucket and the transfer is free

    --- experiments 5 and 6: block and file storage

                        BLOCK (EBS)               FILE (EFS)                OBJECT (S3)
      looks like        a raw disk                a mounted filesystem      an HTTP API
      attached to       ONE instance*             MANY instances            anything, anywhere
      access unit       a 512 B block             a file, byte ranges       a whole object
      in-place edit     yes                       yes                       NO -- rewrite it
      directories       whatever the FS does      real                      NONE -- prefixes
      latency           sub-millisecond           low millisecond           tens of ms
      capacity          provisioned, fixed        elastic                   unlimited
      survives instance yes, if not root          yes                       yes
      $/GB-month        0.080                     0.30                      0.023
      * EBS Multi-Attach exists for io1/io2 and needs a cluster filesystem

    1 TB for a month, by type:
      EBS gp3             $81.92
      EFS Standard       $307.20
      S3 Standard         $23.55
         EFS costs 13x what S3 does and 3.8x what EBS does,
         and it is worth it precisely when several instances must
         share a POSIX filesystem. Paying 13x for a dataset that one
         batch job reads once is the mistake -- that dataset belongs
         in S3.
         Choose by ACCESS PATTERN, not by price: block for a database
         or a boot disk, file for shared POSIX, object for everything
         a data pipeline reads

    provisioned against consumed, on a 1 TB EBS volume holding 200 GB:
      EBS billed on PROVISIONED size : $81.92/month
      S3 billed on STORED bytes      : $4.60/month
         a factor of 18. EBS bills the volume you asked for,
         empty or not; S3 bills the bytes you actually stored. An
         over-provisioned volume is invisible waste, which is why
         'just make it 1 TB to be safe' is an expensive habit

THERE IS NO RENAME

'rename' README.md -> docs/README.md
  bytes read 1,024, bytes written 1,024, API calls 2

Copy plus delete. Renaming a 5 TB dataset "to tidy the folders" moves 10 TB and is billed for it.

VERSIONING BILLS FOR EVERY VERSION

versioning ON, overwrite, then delete:
  older versions kept : 2
  current object      : DeleteMarker

A delete writes a marker; the data is still there and still billed. Set the lifecycle rule when you enable versioning, not later.

Storage classes — 1 TB for a year, retrieved once:

Class Storage/yr Retrieve Total Min days
Standard $282.62 $0.00 $282.62 0
Standard-IA $153.60 $10.24 $163.84 30
Glacier Instant $49.15 $30.72 $79.87 90
Deep Archive $12.17 $20.48 $32.65 180

Read the two ratios separately. Deep Archive storage is 23× cheaper. Add one retrieval a year and the all-in saving falls to 8.7×, because the retrieval fee ($20.48) exceeds a whole year of its storage ($12.17). The headline discount is not the discount.

AND THE REVERSAL

The same 1 TB, retrieved twice a month:

Class Total/yr
Standard $282.62
Standard-IA $399.36
Glacier Instant $786.43

Standard is now the cheapest. "We moved everything to IA to save money" is how a bill goes UP.

Egress:

Transfer Cost
1 TB in $0.00
1 TB out $92.16
1 TB S3 → EC2, same region $0.00

Downloading 1 TB once costs as much as storing it for 3.9 months. Ingress is free; egress is not — and that is the mechanism behind lock-in: your data is not held hostage, it is simply expensive to move.

RESULT

A listing by prefix finds no objects at raw/, only a common prefix; a rename is a copy and a delete; Deep Archive is 23× cheaper to store and 8.7× cheaper all-in.

Experiment 5 — Block storage

1. Question

Launch an instance and configure block storage (EBS).

2. Aim

Attach, format and mount a volume, and price block storage against file and object storage.

3. Steps

The procedure, on the console and CLI, 05_ebs.md:

  1. Launch, attach, format and mount.
  2. Avoid the three mistakes.
  3. Keep it in one zone.
  4. Choose a volume type.
  5. Snapshot it.

BLOCK, FILE AND OBJECT

EBS gp3 EFS Standard S3 Standard
1 TB/month $81.92 $307.20 $23.55

And provisioned against consumed: a 1 TB EBS volume holding 200 GB bills $81.92 where S3 bills $4.60 — a factor of 18. "Just make it 1 TB to be safe" is an expensive habit.

4. Programme

The procedure, on the console and CLI, 05_ebs.md:

# Experiment 5 -- launch an instance and configure block storage (EBS)

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`04_storage.py`, which compares block, file and object semantics and cost**.

---

<!-- Step 1: Launch, attach, format and mount -->
## Launch, attach, format, mount

```bash
aws ec2 run-instances --image-id ami-xxxx --instance-type t3.micro \
  --key-name mykey --security-group-ids sg-xxxx --subnet-id subnet-xxxx

aws ec2 create-volume --size 20 --volume-type gp3 \
  --availability-zone us-east-1a
aws ec2 attach-volume --volume-id vol-xxxx --instance-id i-xxxx \
  --device /dev/sdf
```

Then, on the instance:

```bash
lsblk                                   # find it -- it appears as /dev/nvme1n1
sudo file -s /dev/nvme1n1               # "data" means NO FILESYSTEM yet
sudo mkfs -t xfs /dev/nvme1n1           # DESTROYS anything on it
sudo mkdir /data && sudo mount /dev/nvme1n1 /data
sudo blkid                              # get the UUID
echo 'UUID=<uuid> /data xfs defaults,nofail 0 2' | sudo tee -a /etc/fstab
```

<!-- Step 2: Avoid the three mistakes -->
## The three that go wrong

**1. `mkfs` on the wrong device.** `lsblk` first, every time. Formatting the
root volume ends the instance.

**2. Forgetting `/etc/fstab`.** The volume is not mounted after a reboot and
your application starts writing to the root disk instead — silently, until it
fills.

**3. `nofail` omitted.** If the volume is missing at boot, the instance hangs
in the boot sequence and you cannot SSH in to fix it. `nofail` turns a
catastrophe into a missing directory.

<!-- Step 3: Keep it in one zone -->
## An EBS volume is in ONE availability zone

You cannot attach `us-east-1a`'s volume to an instance in `us-east-1b`. To
move it: snapshot it (snapshots go to S3, which is regional), then create a
volume from the snapshot in the other AZ.

**That constraint is why "an EBS volume attaches to one instance" is really
"one instance, in one AZ"** — and it is why EFS exists.

<!-- Step 4: Choose a volume type -->
## Volume types, and how to choose

| Type | For | Notes |
|---|---|---|
| **gp3** | almost everything | IOPS and throughput are **independent of size** |
| gp2 | legacy | IOPS scale with size — the reason gp3 replaced it |
| io2 Block Express | a demanding database | expensive, and supports Multi-Attach |
| st1 | big sequential reads | HDD; terrible at random I/O |
| sc1 | cold archive on a disk | cheapest, slowest |

**gp3 is the default answer.** Under gp2 you would over-provision a volume
purely to buy IOPS; gp3 unbundled them.

<!-- Step 5: Snapshot it -->
## Snapshots

```bash
aws ec2 create-snapshot --volume-id vol-xxxx --description "before upgrade"
```

Snapshots are **incremental** — only changed blocks — and stored in S3.
Deleting an intermediate snapshot does not break later ones; AWS re-parents
the blocks. But snapshots of a **mounted, busy filesystem** may be
crash-consistent rather than clean: freeze or unmount for a database.

5. Execution and Results

The procedure, on the console and CLI, 05_ebs.md:

NOT RUN HERE

05_ebs.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model for this experiment, 04_storage.py, is shown in full under Experiment 4, with what it printed.

RESULT

1 TB costs $81.92 a month on EBS; a 1 TB volume holding 200 GB bills $81.92 where S3 would bill $4.60.

Experiment 6 — File storage

1. Question

Create and configure file storage on a cloud VM (EFS).

2. Aim

Create and mount a shared file system, and decide when it is worth its price.

3. Steps

The procedure, on the console and CLI, 06_efs.md:

  1. Create and mount.
  2. Open the security group.
  3. Decide what EFS is for.
  4. Count the cost.
  5. Know the equivalents.

THE PRICE

EFS costs 13× S3 and 3.8× EBS — worth it precisely when several instances must share a POSIX filesystem, and a mistake for a dataset one batch job reads once. The figures are the table under Experiment 5.

4. Programme

The procedure, on the console and CLI, 06_efs.md:

# Experiment 6 -- create and configure file storage on a cloud VM (EFS)

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`04_storage.py`, which prices EFS against EBS and S3 on the same terabyte**.

---

<!-- Step 1: Create and mount -->
## Create and mount

```bash
aws efs create-file-system --performance-mode generalPurpose \
    --throughput-mode elastic --tags Key=Name,Value=shared-data
aws efs create-mount-target --file-system-id fs-xxxx \
    --subnet-id subnet-xxxx --security-groups sg-xxxx
```

On each instance:

```bash
sudo apt install -y amazon-efs-utils
sudo mkdir /shared
sudo mount -t efs -o tls fs-xxxx:/ /shared
echo 'fs-xxxx:/ /shared efs _netdev,tls 0 0' | sudo tee -a /etc/fstab
```

**`_netdev` is not optional.** It tells systemd to wait for the network
before mounting. Without it the mount fails at boot, every time.

<!-- Step 2: Open the security group -->
## The security group rule everyone forgets

The EFS mount target needs an inbound rule allowing **NFS (TCP 2049)** from
the instances' security group. Without it the mount hangs — it does not fail
with a useful message, it hangs — and this is the single most common EFS
problem.

<!-- Step 3: Decide what EFS is for -->
## What EFS is actually for

Mount it on **many instances at once**, and they all see the same files with
POSIX semantics — locking, permissions, in-place writes. That is the whole
feature, and nothing else in the storage lineup offers it.

| Use it for | Do not use it for |
|---|---|
| a shared home directory across a fleet | a dataset one batch job reads once |
| shared model artefacts several workers read | a database's data files |
| a lift-and-shift app that expects a filesystem | anything a pipeline can read from S3 |

<!-- Step 4: Count the cost -->
## Cost, which is the reason to think twice

EFS Standard is about **$0.30/GB-month** — roughly 13x S3 Standard and
nearly 4x EBS gp3. Use **EFS Infrequent Access** with a lifecycle policy
(`--lifecycle-policy TransitionToIA=AFTER_30_DAYS`) and the cold portion
drops by about 90%.

**"We put the training data on EFS because it was easy to mount" is how a
storage bill triples.** That data belongs in S3.

<!-- Step 5: Know the equivalents -->
## Azure and GCP

| | AWS | Azure | GCP |
|---|---|---|---|
| Managed NFS | EFS | Azure Files (NFS/SMB) | Filestore |
| Protocol | NFSv4.1 | SMB **and** NFS | NFSv3 |

**Azure Files speaks SMB**, which matters if your workload is Windows —
that is the one real difference between the three.

5. Execution and Results

The procedure, on the console and CLI, 06_efs.md:

NOT RUN HERE

06_efs.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model for this experiment, 04_storage.py, is shown in full under Experiment 4, with what it printed.

RESULT

EFS costs 13× S3 and 3.8× EBS: worth it when several instances share a POSIX file system.

Experiment 7 — The notebook environment

1. Question

Set up Jupyter Notebook or Colab on a cloud VM.

2. Aim

Run a notebook's cells in order and assert their outputs, and never publish it unauthenticated.

3. Steps

The procedure, on the console and CLI, 07_notebook.md:

  1. Start the notebook on the VM.
  2. Avoid the mistake that matters.
  3. Or use a managed notebook.
  4. Set the lifecycle configuration.
  5. Compare it with Colab.

WHAT RUNS

Cells executed and asserted in 01_vm_and_hosting.py:

In  [1]: import pandas as pd; import fixtures as f
In  [2]: df.shape                       -> (9, 19)
In  [3]: df.groupby('region')['revenue'].sum()  -> South 10360.0
In  [4]: df['revenue'].sum()            -> 12880.0

Four cells, executed in order, every output asserted. That is what a notebook test looks like — papermill and nbconvert --execute do exactly this in CI, and a notebook nobody executes in CI is a notebook that has already drifted.

4. Programme

The procedure, on the console and CLI, 07_notebook.md:

# Experiment 7 -- set up Jupyter Notebook / Colab on a cloud VM

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`01_vm_and_hosting.py`, which executes notebook cells and asserts every output**.

---

<!-- Step 1: Start the notebook on the VM -->
## On the VM

```bash
sudo apt install -y python3-pip
pip install jupyterlab pandas scikit-learn matplotlib
jupyter lab --generate-config
jupyter lab password                      # set one, do not skip this
```

Then, **do not** open it to the internet. Use an SSH tunnel:

```bash
# on the VM
jupyter lab --no-browser --port=8888 --ip=127.0.0.1

# on your machine
ssh -N -L 8888:localhost:8888 ubuntu@<vm-ip>
# then browse to http://localhost:8888
```

<!-- Step 2: Avoid the mistake that matters -->
## ⚠ The mistake that matters

```bash
jupyter lab --ip=0.0.0.0 --allow-root --NotebookApp.token=''
```

**That publishes a root shell on the internet.** A Jupyter notebook executes
arbitrary code by design, so an unauthenticated notebook is not "an insecure
notebook" — it is a remote code execution endpoint. Scanners find these in
minutes; it is a standard way cloud accounts get used for cryptomining.

**Always: an SSH tunnel, or a managed notebook behind IAM.**

<!-- Step 3: Or use a managed notebook -->
## SageMaker Studio / Vertex Workbench instead

```bash
aws sagemaker create-notebook-instance \
  --notebook-instance-name lab7 --instance-type ml.t3.medium \
  --role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole
```

You get the tunnel, the authentication and the IAM role for free — **and no
access key is ever written to disk**, which is the real argument for it.

<!-- Step 4: Set the lifecycle configuration -->
## The lifecycle configuration that saves the money

```bash
#!/bin/bash
# attach as an OnStart lifecycle config
IDLE_TIME=3600
pip install -q jupyter-autoshutdown-extension || true
echo "auto-shutdown after ${IDLE_TIME}s idle" >> /var/log/lifecycle.log
```

**Colab disconnects after about 90 minutes idle and that is an annoyance. A
cloud notebook does not disconnect, and that is a bill.** An idle-shutdown
policy is the single most useful thing to configure on day one.

<!-- Step 5: Compare it with Colab -->
## Colab against a cloud notebook

| | Colab | Cloud notebook |
|---|---|---|
| Cost | free tier, then paid | the **instance**, hourly |
| Data | upload, or mount Drive | IAM role, no keys |
| Stops when | idle ~90 min | **never — you stop it** |
| State on stop | **lost** | kept on the volume |
| GPU | when available | the one you pay for |
| Private data | a policy question | inside your VPC |

**Colab is excellent for learning and wrong for anything confidential.** The
deciding question is whose infrastructure the data may sit on.

5. Execution and Results

The procedure, on the console and CLI, 07_notebook.md:

NOT RUN HERE

07_notebook.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

THE MISTAKE THAT MATTERS

jupyter lab --ip=0.0.0.0 --allow-root --NotebookApp.token=''

That publishes a root shell on the internet. A notebook executes arbitrary code by design, so an unauthenticated one is not "an insecure notebook" — it is a remote code execution endpoint. Scanners find these in minutes.

Always an SSH tunnel, or a managed notebook behind IAM.

AND THE ROW THAT COSTS MONEY

Colab Cloud notebook
Stops when idle ~90 min never — you stop it
State on stop lost kept on the volume

An m5.xlarge notebook left running costs about $140/month. Colab disconnecting is an annoyance; a cloud notebook not disconnecting is a bill. Set an idle-shutdown lifecycle policy on day one.

The Python model for this experiment, 01_vm_and_hosting.py, is shown in full under Experiment 1, with what it printed.

RESULT

Four cells executed in order, every output asserted: South ₹10,360 and ₹12,880 in all.

Experiment 8 — Cloud-hosted databases

1. Question

Connect to cloud-hosted database services: RDS, BigQuery and Cosmos DB.

2. Aim

Query a managed database and a serverless warehouse, and see what a query costs.

3. Steps

The procedure, on the console and CLI, 08_cloud_db.md:

  1. Create a managed database.
  2. Query BigQuery.
  3. Try Cosmos DB.
  4. Compare managed with self-hosted.

WHAT A QUERY COSTS

BigQuery, at $6.25 per TB scanned:

Query TB scanned Cost
SELECT * FROM events 10.00 $62.50
SELECT user_id FROM events 0.40 $2.50
SELECT user_id … WHERE dt = '…' 0.02 $0.12

The same question, 500× the price. Column projection and partition pruning — Big Data Technologies' techniques, saving money here instead of time. That is why SELECT * is a billing incident on a serverless warehouse and merely rude on a server you already own.

4. Programme

The procedure, on the console and CLI, 08_cloud_db.md:

# Experiment 8 -- connect to cloud-hosted database services (RDS, BigQuery, Cosmos DB)

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`09_etl_warehouse.py`, which runs the same extract and the same warehouse queries locally**.

---

<!-- Step 1: Create a managed database -->
## RDS (managed PostgreSQL/MySQL)

```bash
aws rds create-db-instance \
  --db-instance-identifier retail-db --db-instance-class db.t3.micro \
  --engine postgres --allocated-storage 20 \
  --master-username admin --manage-master-user-password \
  --no-publicly-accessible --backup-retention-period 7
```

```bash
psql -h retail-db.xxxx.us-east-1.rds.amazonaws.com -U admin -d postgres
```

**`--no-publicly-accessible` is the right default.** Reach it from an EC2
instance in the same VPC, or through a bastion, or with SSM Session Manager.
A publicly reachable database with a weak password is found by scanners in
hours.

**And the security group needs an inbound rule for port 5432 from your
instance's security group** — group-to-group, not a CIDR. That is the fix for
almost every "connection timed out".

<!-- Step 2: Query BigQuery -->
## BigQuery

```bash
bq mk --dataset retail
bq load --autodetect --source_format=CSV retail.sales gs://bucket/sales.csv

bq query --use_legacy_sql=false --dry_run \
  'SELECT region, SUM(revenue) FROM retail.sales GROUP BY region'
```

**`--dry_run` prints the bytes the query would scan without running it, and
therefore what it would cost.** Run it before every large query. It is the
single most valuable BigQuery habit.

```bash
bq query --use_legacy_sql=false --maximum_bytes_billed=1000000000 '...'
```

**`--maximum_bytes_billed` is a seatbelt**: the query fails rather than
billing more than you said. Set it in every automated job.

<!-- Step 3: Try Cosmos DB -->
## Cosmos DB

```bash
az cosmosdb create --name retail-cosmos --resource-group rg \
  --kind GlobalDocumentDB --default-consistency-level Session
az cosmosdb sql database create --account-name retail-cosmos \
  --resource-group rg --name retail
```

**The consistency level is the interesting choice**, and it is Course 10's
CAP discussion made into a dropdown:

| Level | Guarantee | Cost |
|---|---|---|
| Strong | linearizable | highest RU, single region write |
| **Bounded staleness** | lag bounded by time or versions | high |
| **Session** | your own writes are read back | **the default, and usually right** |
| Consistent prefix | order preserved, some lag | low |
| Eventual | none | **lowest RU** |

**Cosmos bills in Request Units**, and stronger consistency costs more RUs
per read. That is CAP with a price tag attached.

<!-- Step 4: Compare managed with self-hosted -->
## Managed against self-hosted

| | Self-hosted on EC2 | Managed (RDS) |
|---|---|---|
| Patching | you | automatic, in a window |
| Backups | you write them | continuous, point-in-time |
| Failover | you build it | Multi-AZ, ~60 s |
| Version upgrade | your weekend | a click, with a restart |
| Root/superuser | yes | **no** |
| Cost | instance only | instance + ~20-30% |

**"No superuser" is the row that surprises people.** Some extensions, some
`ALTER SYSTEM` settings and any filesystem access are unavailable. If your
application needs those, managed is not an option — and knowing that before
you migrate is the point.

5. Execution and Results

The procedure, on the console and CLI, 08_cloud_db.md:

NOT RUN HERE

08_cloud_db.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model for this experiment, 09_etl_warehouse.py, is shown in full under Experiment 9, with what it printed.

RESULT

The same question costs $62.50 as SELECT * and $0.12 with a column list and a partition filter.

Experiment 9 — A batch ETL pipeline

1. Question

Build a batch ETL pipeline: extract from an operational database, transform, and load a warehouse.

2. Aim

Run the pipeline end to end, with an audit trail at every step, and check South against three other engines.

3. Steps

The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:

  1. Extract from the operational database.
  2. Transform, with an audit trail.
  3. Load into the warehouse.
  4. Set ETL against ELT.
  5. See what makes a cloud warehouse different.
  6. Price the queries.

THE PIPELINE, WITH AN AUDIT TRAIL

Step Rows
extracted 11
after dedup 10
dropped, null region 1
loaded 9

11 in, 9 out, and the pipeline can say where the other two went. A transformation that silently drops rows is worse than one that fails: the numbers still look plausible. Every ETL job should emit these counts, and a monitoring rule should alarm when the drop rate moves.

4. Programme

The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:

"""Experiments 8, 9 and 12 -- cloud databases, a batch ETL pipeline, and
loading a cloud data warehouse.

AWS RDS, BigQuery and Cosmos DB need an account, so `08_cloud_db.md` and
`12_etl_to_warehouse.md` carry the console and CLI steps, marked NOT
EXECUTED.

What runs is the pipeline itself: extract from a real relational source
(SQLite, standing in for RDS), transform, and load into a real columnar
warehouse (DuckDB, standing in for Redshift/BigQuery). The transformations,
the row counts and the money all reconcile -- and they reconcile against
Course 11 and Course 12 B, which used the same nine facts.

The pricing arithmetic at the end is the part of "cloud data warehouse" that
is genuinely different from Course 5's DBMS, and it is what gets examined.
"""
import os
import sqlite3
import tempfile

import duckdb

import fixtures as f

SRC_COLUMNS = ["order_id", "date_key", "store", "region", "product",
               "category", "qty", "list_price", "unit_cost"]


def money(x):
    return f"${x:,.2f}"


def extract(db_path):
    """E -- read from the operational database. Never transform here."""
    con = sqlite3.connect(db_path)
    con.execute("""CREATE TABLE orders (
        order_id INTEGER PRIMARY KEY, date_key TEXT, store TEXT, region TEXT,
        product TEXT, category TEXT, qty INTEGER,
        list_price REAL, unit_cost REAL)""")
    rows = []
    for i, (_, r) in enumerate(f.SALES_DF.iterrows()):
        rows.append((i + 1, r["date_key"], r["store"], r["region"],
                     r["product"], r["category"], int(r["qty"]),
                     float(r["list_price"]), float(r["unit_cost"])))
    # deliberately dirty: one duplicate and one null region
    rows.append((10, rows[0][1], rows[0][2], rows[0][3], rows[0][4],
                 rows[0][5], rows[0][6], rows[0][7], rows[0][8]))
    rows.append((11, "D4", "Guntur", None, "Tea 500g", "Grocery",
                 3, 210.0, 150.0))
    con.executemany(
        "INSERT INTO orders VALUES (?,?,?,?,?,?,?,?,?)", rows)
    con.commit()
    out = con.execute(
        f"SELECT {', '.join(SRC_COLUMNS)} FROM orders").fetchall()
    con.close()
    return out


def transform(rows, log):
    """T -- and every step records what it dropped, or it is not auditable."""
    log["extracted"] = len(rows)

    seen, deduped = set(), []
    for r in rows:
        fingerprint = r[1:]                 # everything but the surrogate id
        if fingerprint in seen:
            continue
        seen.add(fingerprint)
        deduped.append(r)
    log["after_dedup"] = len(deduped)

    clean = [r for r in deduped if r[3] is not None]
    log["dropped_null_region"] = len(deduped) - len(clean)

    enriched = []
    for r in clean:
        d = dict(zip(SRC_COLUMNS, r))
        d["revenue"] = d["qty"] * d["list_price"]
        d["cost"] = d["qty"] * d["unit_cost"]
        d["profit"] = d["revenue"] - d["cost"]
        d["quarter"] = "Q1" if d["date_key"] in ("D1", "D2") else "Q2"
        enriched.append(d)
    log["loaded"] = len(enriched)
    return enriched


def main():
    print("  Experiments 8, 9 and 12 -- cloud DB, batch ETL, warehouse load")

    tmp = tempfile.mkdtemp(prefix="cloud09_")
    db = os.path.join(tmp, "orders.db")

    # Step 1: Extract from the operational database
    raw = extract(db)
    print(f"\n    EXTRACT from the operational database (SQLite as RDS):")
    print(f"      {len(raw)} rows, including one duplicate and one null region")
    assert len(raw) == 11

    # Step 2: Transform, with an audit trail
    log = {}
    clean = transform(raw, log)
    print(f"\n    TRANSFORM, with an audit trail at every step:")
    print(f"      {'step':<26}{'rows':>6}")
    for step in ("extracted", "after_dedup", "dropped_null_region", "loaded"):
        print(f"      {step:<26}{log[step]:>6}")
    assert log == {"extracted": 11, "after_dedup": 10,
                   "dropped_null_region": 1, "loaded": 9}
    print("""         11 in, 9 out, and the pipeline can SAY WHERE THE OTHER TWO
         WENT. A transformation that silently drops rows is worse than
         one that fails: the numbers still look plausible.
         Every ETL job should emit these counts, and a monitoring rule
         should alarm when the drop rate moves""")

    # Step 3: Load into the warehouse
    con = duckdb.connect()
    con.execute("""CREATE TABLE fact_sales (
        order_id INTEGER, date_key VARCHAR, store VARCHAR, region VARCHAR,
        product VARCHAR, category VARCHAR, qty INTEGER,
        list_price DOUBLE, unit_cost DOUBLE,
        revenue DOUBLE, cost DOUBLE, profit DOUBLE, quarter VARCHAR)""")
    con.executemany(
        "INSERT INTO fact_sales VALUES (" + ",".join(["?"] * 13) + ")",
        [[r[c] for c in ("order_id", "date_key", "store", "region", "product",
                         "category", "qty", "list_price", "unit_cost",
                         "revenue", "cost", "profit", "quarter")]
         for r in clean])

    n, rev = con.execute(
        "SELECT COUNT(*), SUM(revenue) FROM fact_sales").fetchone()
    print(f"\n    LOAD into the warehouse (DuckDB as Redshift/BigQuery):")
    print(f"      {n} rows, revenue {money(rev)}")
    assert n == 9 and rev == f.total_revenue()
    print(f"""         {money(rev)} is Course 11's total, Course 12 B's Hive total and
         Course 12 B's Spark total. FOUR engines now agree on the same
         nine facts, and the suite fails if any of them drifts""")

    rows = con.execute("""
        SELECT region, SUM(revenue) rev, SUM(profit) prof
        FROM fact_sales GROUP BY region ORDER BY rev DESC""").fetchall()
    print(f"\n      {'region':<10}{'revenue':>12}{'profit':>10}")
    for region, r, p in rows:
        print(f"      {region:<10}{money(r):>12}{money(p):>10}")
    assert dict((r[0], r[1]) for r in rows)["South"] == 10360.0

    # Step 4: Set ETL against ELT
    print("\n    ETL against ELT, which is the modern distinction:")
    print(f"      {'':<16}{'ETL':<34}{'ELT'}")
    for label, etl, elt in (
            ("transform runs", "on a separate compute box", "IN the warehouse"),
            ("lands in the DW", "clean data only", "RAW data, then transformed"),
            ("re-run a change", "re-extract from source", "re-run SQL on raw"),
            ("needs", "an ETL server or Glue", "a warehouse that scales"),
            ("source load", "one read", "one read"),
            ("suits", "limited warehouse capacity", "cheap elastic compute")):
        print(f"      {label:<16}{etl:<34}{elt}")
    print("""         ELT WON BECAUSE WAREHOUSE COMPUTE GOT CHEAP AND ELASTIC.
         Landing raw data means a transformation bug is fixed by
         re-running SQL rather than re-extracting from a production
         database that may no longer hold the old rows -- which is
         exactly the DELETE problem Course 12 B found in Sqoop""")

    # Step 5: See what makes a cloud warehouse different
    print("\n    what makes a cloud DW different from Course 5's RDBMS:")
    print(f"      {'':<22}{'RDBMS (Course 5)':<28}{'cloud DW'}")
    for label, rdbms, dw in (
            ("storage layout", "ROW", "COLUMNAR"),
            ("workload", "many small transactions", "few huge scans"),
            ("indexes", "central to performance", "usually none"),
            ("scaling", "a bigger box", "add nodes / serverless"),
            ("compute & storage", "coupled", "SEPARATED"),
            ("billed on", "the box, hourly", "BYTES SCANNED or node-hours"),
            ("a bad query costs", "time", "TIME AND MONEY")):
        print(f"      {label:<22}{rdbms:<28}{dw}")

    # Step 6: Price the queries
    print("\n    BigQuery on-demand, at "
          f"${f.BIGQUERY_PER_TB_SCANNED:.2f} per TB SCANNED:")
    print(f"      {'query':<44}{'TB scanned':>12}{'cost':>10}")
    scenarios = [
        ("SELECT * FROM events                      ", 10.0),
        ("SELECT user_id FROM events                ", 0.4),
        ("SELECT user_id ... WHERE dt = '2026-08-01'", 0.02),
    ]
    costs = {}
    for label, tb in scenarios:
        c = tb * f.BIGQUERY_PER_TB_SCANNED
        costs[label.strip()] = c
        print(f"      {label:<44}{tb:>12.2f}{money(c):>10}")
    full = costs["SELECT * FROM events"]
    pruned = costs["SELECT user_id ... WHERE dt = '2026-08-01'"]
    assert full / pruned == 500
    print(f"""         THE SAME QUESTION, {full / pruned:.0f}x THE PRICE. Selecting one
         column instead of all reads a fraction of the bytes
         (Course 12 B's column projection), and a partition filter
         removes almost all of the rest (Course 12 B's partition
         pruning).
         In Course 12 B those techniques saved TIME. Here they save
         MONEY, on the same mechanism -- which is why 'SELECT *' is a
         billing incident on a serverless warehouse and merely rude on
         a server you already own""")

    print(f"\n    Redshift, at ${f.REDSHIFT_RA3_XLPLUS_HOUR:.3f} per node-hour:")
    for nodes in (2, 4, 8):
        m = nodes * f.REDSHIFT_RA3_XLPLUS_HOUR * f.HOURS_PER_MONTH
        print(f"      {nodes} nodes: {money(m):>12}/month, "
              f"{'queries are free at the margin'}")
    two = 2 * f.REDSHIFT_RA3_XLPLUS_HOUR * f.HOURS_PER_MONTH
    breakeven_tb = two / f.BIGQUERY_PER_TB_SCANNED
    print(f"""
      break-even: {breakeven_tb:,.0f} TB scanned per month
         BELOW that, on-demand BigQuery is cheaper and you pay nothing
         when idle. ABOVE it, a provisioned cluster is cheaper and an
         extra query costs nothing at the margin.
         That is the whole provisioned-against-serverless decision,
         and it is a calculation rather than a preference""")
    assert 250 < breakeven_tb < 260

    con.close()
    os.remove(db)
    os.rmdir(tmp)


if __name__ == "__main__":
    main()

5. Execution and Results

The Python model, which runs, 09_etl_warehouse.py, for experiments 8, 9 and 12:

OUTPUT

  Experiments 8, 9 and 12 -- cloud DB, batch ETL, warehouse load

    EXTRACT from the operational database (SQLite as RDS):
      11 rows, including one duplicate and one null region

    TRANSFORM, with an audit trail at every step:
      step                        rows
      extracted                     11
      after_dedup                   10
      dropped_null_region            1
      loaded                         9
         11 in, 9 out, and the pipeline can SAY WHERE THE OTHER TWO
         WENT. A transformation that silently drops rows is worse than
         one that fails: the numbers still look plausible.
         Every ETL job should emit these counts, and a monitoring rule
         should alarm when the drop rate moves

    LOAD into the warehouse (DuckDB as Redshift/BigQuery):
      9 rows, revenue $12,880.00
         $12,880.00 is Course 11's total, Course 12 B's Hive total and
         Course 12 B's Spark total. FOUR engines now agree on the same
         nine facts, and the suite fails if any of them drifts

      region         revenue    profit
      South       $10,360.00 $2,760.00
      North        $2,520.00   $765.00

    ETL against ELT, which is the modern distinction:
                      ETL                               ELT
      transform runs  on a separate compute box         IN the warehouse
      lands in the DW clean data only                   RAW data, then transformed
      re-run a change re-extract from source            re-run SQL on raw
      needs           an ETL server or Glue             a warehouse that scales
      source load     one read                          one read
      suits           limited warehouse capacity        cheap elastic compute
         ELT WON BECAUSE WAREHOUSE COMPUTE GOT CHEAP AND ELASTIC.
         Landing raw data means a transformation bug is fixed by
         re-running SQL rather than re-extracting from a production
         database that may no longer hold the old rows -- which is
         exactly the DELETE problem Course 12 B found in Sqoop

    what makes a cloud DW different from Course 5's RDBMS:
                            RDBMS (Course 5)            cloud DW
      storage layout        ROW                         COLUMNAR
      workload              many small transactions     few huge scans
      indexes               central to performance      usually none
      scaling               a bigger box                add nodes / serverless
      compute & storage     coupled                     SEPARATED
      billed on             the box, hourly             BYTES SCANNED or node-hours
      a bad query costs     time                        TIME AND MONEY

    BigQuery on-demand, at $6.25 per TB SCANNED:
      query                                         TB scanned      cost
      SELECT * FROM events                               10.00    $62.50
      SELECT user_id FROM events                          0.40     $2.50
      SELECT user_id ... WHERE dt = '2026-08-01'          0.02     $0.12
         THE SAME QUESTION, 500x THE PRICE. Selecting one
         column instead of all reads a fraction of the bytes
         (Course 12 B's column projection), and a partition filter
         removes almost all of the rest (Course 12 B's partition
         pruning).
         In Course 12 B those techniques saved TIME. Here they save
         MONEY, on the same mechanism -- which is why 'SELECT *' is a
         billing incident on a serverless warehouse and merely rude on
         a server you already own

    Redshift, at $1.086 per node-hour:
      2 nodes:    $1,585.56/month, queries are free at the margin
      4 nodes:    $3,171.12/month, queries are free at the margin
      8 nodes:    $6,342.24/month, queries are free at the margin

      break-even: 254 TB scanned per month
         BELOW that, on-demand BigQuery is cheaper and you pay nothing
         when idle. ABOVE it, a provisioned cluster is cheaper and an
         extra query costs nothing at the margin.
         That is the whole provisioned-against-serverless decision,
         and it is a calculation rather than a preference

This experiment has no console procedure: SQLite stands in for the operational database (RDS) and DuckDB for the warehouse, and the pipeline between them runs end to end. Experiment 12's procedure does the same with Glue and Redshift or BigQuery.

THE FOUR-ENGINE CHECK

Region Revenue Profit Margin
South 10,360 2,760 26.64%
North 2,520 765 30.36%

₹12,880 total, ₹10,360 for South — Business Intelligence Tools' DAX, Big Data Technologies' Hive, Big Data Technologies' Spark and this. Four engines, one set of nine facts.

And the break-even: Redshift at $1.086/node-hour: 2 nodes is $1,585.56/month. Break-even against on-demand BigQuery: about 254 TB scanned per month. Below that, serverless is cheaper and costs nothing when idle. Above it, a cluster is cheaper and an extra query is free at the margin. A calculation, not a preference.

RESULT

11 rows in, 9 loaded, and the pipeline says where the other two went; South is ₹10,360, as in DAX, Hive and Spark.

Experiment 10 — A SageMaker notebook with an IAM role

1. Question

Launch a SageMaker notebook, and attach an IAM role and an S3 bucket.

2. Aim

Give the notebook a role rather than keys, and know what an idle endpoint costs.

3. Steps

The procedure, on the console and CLI, 10_sagemaker_notebook.md:

  1. Create the role first.
  2. Then the notebook.
  3. Avoid what not to do.
  4. Stop it when you are done.

A ROLE IS NOT A USER

User Role
Credentials long-lived access key temporary, auto-rotated
In a notebook keys in a file — bad attached; no keys exist
If leaked valid until revoked expires in minutes to hours

A SageMaker notebook gets an execution role, so no access key is ever written to disk. That is why experiment 10 says "attach IAM role" rather than "paste your credentials", and "I put my keys in the notebook" is the answer that loses the marks.

4. Programme

The procedure, on the console and CLI, 10_sagemaker_notebook.md:

# Experiment 10 -- launch a SageMaker notebook, attach an IAM role and an S3 bucket

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`03_iam_and_account.py`, which implements IAM's evaluation algorithm and proves what the role can do**.

---

<!-- Step 1: Create the role first -->
## The role first, then the notebook

```bash
aws iam create-role --role-name SageMakerExecutionRole \
  --assume-role-policy-document '{
    "Version": "2012-10-17",
    "Statement": [{
      "Effect": "Allow",
      "Principal": {"Service": "sagemaker.amazonaws.com"},
      "Action": "sts:AssumeRole"}]}'
```

**That document is the TRUST POLICY, and it is not the permissions.** It says
*who may become this role*; a separate permissions policy says *what the role
may then do*. Confusing the two is the commonest IAM error after the explicit
Deny.

```bash
aws iam put-role-policy --role-name SageMakerExecutionRole \
  --policy-name S3Scoped --policy-document '{
    "Version": "2012-10-17",
    "Statement": [
      {"Effect": "Allow", "Action": ["s3:GetObject"],
       "Resource": "arn:aws:s3:::retail-lake/train/*"},
      {"Effect": "Allow", "Action": ["s3:PutObject"],
       "Resource": "arn:aws:s3:::retail-lake/models/*"}]}'
```

**Two statements, two prefixes, two actions.** Not `s3:*` on `*`. The
runnable half shows both policies letting the training job succeed, and only
one of them also permitting `iam:CreateUser`.

<!-- Step 2: Then the notebook -->
## Then the notebook

```bash
aws sagemaker create-notebook-instance \
  --notebook-instance-name lab10 \
  --instance-type ml.t3.medium \
  --role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole \
  --volume-size-in-gb 20
aws sagemaker describe-notebook-instance --notebook-instance-name lab10
aws sagemaker create-presigned-notebook-instance-url \
  --notebook-instance-name lab10
```

Inside the notebook, **no credentials exist anywhere**:

```python
import boto3, sagemaker
session = sagemaker.Session()
role = sagemaker.get_execution_role()      # reads the ATTACHED role
bucket = "retail-lake"
session.upload_data("train.csv", bucket=bucket, key_prefix="train")
```

<!-- Step 3: Avoid what not to do -->
## ⚠ What not to do

```python
boto3.client("s3", aws_access_key_id="AKIA...", aws_secret_access_key="...")
```

**Keys in a notebook end up in git.** GitHub scans for them and AWS
quarantines the account, which is the good outcome; the bad outcome is a
cryptomining bill. The execution role exists precisely so this is never
necessary.

<!-- Step 4: Stop it when you are done -->
## Stop it when you are done

```bash
aws sagemaker stop-notebook-instance --notebook-instance-name lab10
aws sagemaker delete-notebook-instance --notebook-instance-name lab10
```

**Stopping keeps the EBS volume (and its charge); deleting removes
everything.** A stopped notebook costs only storage; a running one costs the
instance whether or not the browser tab is open.

5. Execution and Results

The procedure, on the console and CLI, 10_sagemaker_notebook.md:

NOT RUN HERE

10_sagemaker_notebook.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

THE ENDPOINT TRAP

A forgotten ml.m5.large endpoint costs about $70/month. A training job ends and stops billing; an endpoint runs until you delete it, at hourly rates, whether or not anything calls it.

Set a budget alarm on day one, before anything else.

The Python model for this experiment, 03_iam_and_account.py, is shown in full under Experiment 3, with what it printed.

RESULT

A role's credentials are temporary and never written to disk; a forgotten endpoint costs about $70 a month.

Experiment 11 — Train a model on a managed platform

1. Question

Build a classification or regression model on a managed ML platform.

2. Aim

Train a model, quote the dummy first, save the artefact, and price the instance.

3. Steps

The procedure, on the console and CLI, 11_sagemaker_train.md:

  1. Call the SDK.
  2. See what makes it managed.
  3. Follow the train.py contract.
  4. Train on spot.
  5. Choose the instance.

The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:

  1. Make the data.
  2. Experiment 11: run the training job.
  3. Save the artefact, and reload it.
  4. See what the cloud changes.
  5. Price the instance.
  6. Experiment 14: run AutoML.
  7. List what AutoML cannot do.
  8. Price the search.

QUOTE THE DUMMY FIRST, ALWAYS

Model Accuracy F1 AUC
DummyClassifier 0.8433 0.0000 0.5000
GradientBoosting 0.9467 0.8095 0.9029

94.67% sounds excellent until you see 84.33% for predicting "never churns". The real gain is 10.3 percentage points, and the F1 of 0.8095 against 0.0000 is what shows the model found anything.

Machine Learning's argument, and it does not stop being true because the model trained on somebody else's computer.

4. Programme

The procedure, on the console and CLI, 11_sagemaker_train.md:

# Experiment 11 -- build a classification/regression model on a managed ML platform

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`11_train_and_automl.py`, which trains the same model and prices the instance choice**.

---

<!-- Step 1: Call the SDK -->
## The SDK call

```python
from sagemaker.sklearn.estimator import SKLearn

estimator = SKLearn(
    entry_point="train.py",
    role=role,
    instance_type="ml.m5.xlarge",
    instance_count=1,
    framework_version="1.2-1",
    hyperparameters={"n_estimators": 100, "max_depth": 5},
    output_path=f"s3://{bucket}/models/",
)
estimator.fit({"train": f"s3://{bucket}/train/",
               "validation": f"s3://{bucket}/validation/"})
```

<!-- Step 2: See what makes it managed -->
## The three things that make it a MANAGED job

1. **`entry_point="train.py"`** — your ordinary scikit-learn script, run
   inside a container SageMaker builds. The algorithm is unchanged.
2. **`instance_type`** — billed per second, spun up for the job and destroyed
   after. This is the actual product.
3. **`output_path`** — the artefact lands in S3. Training and serving are
   separate systems joined by one file.

<!-- Step 3: Follow the train.py contract -->
## The `train.py` contract

```python
import argparse, os, joblib, pandas as pd
from sklearn.ensemble import GradientBoostingClassifier

parser = argparse.ArgumentParser()
parser.add_argument("--n_estimators", type=int, default=100)
parser.add_argument("--train", default=os.environ["SM_CHANNEL_TRAIN"])
parser.add_argument("--model-dir", default=os.environ["SM_MODEL_DIR"])
args = parser.parse_args()

df = pd.read_csv(os.path.join(args.train, "train.csv"))
X, y = df.drop(columns=["target"]), df["target"]
model = GradientBoostingClassifier(n_estimators=args.n_estimators).fit(X, y)
joblib.dump(model, os.path.join(args.model_dir, "model.joblib"))
```

**Hyperparameters arrive as command-line arguments; channels arrive as
environment variables; the model must be written to `SM_MODEL_DIR`.** Those
three conventions are the entire interface, and getting `SM_MODEL_DIR` wrong
is why a job "succeeds" and produces no artefact.

<!-- Step 4: Train on spot -->
## Spot training

```python
estimator = SKLearn(..., use_spot_instances=True,
                    max_run=3600, max_wait=7200,
                    checkpoint_s3_uri=f"s3://{bucket}/checkpoints/")
```

**Up to 70% cheaper, and interruptible.** `max_wait` must exceed `max_run` to
leave room for interruptions, and without `checkpoint_s3_uri` an interrupted
job restarts from zero. For a 20-minute job spot is free money; for a 20-hour
job without checkpoints it is a trap.

<!-- Step 5: Choose the instance -->
## Instance choice, which is answered by the algorithm

| Algorithm | Instance | Why |
|---|---|---|
| scikit-learn, XGBoost on tabular | **m5 / c5** | no GPU code path exists |
| Deep learning, training | p3 / p4d / g5 | dense matrix multiplication |
| Deep learning, inference | g4dn / inf1 | cheaper per prediction |
| Anything, if the data fits in RAM | the smallest that fits | you are paying for RAM |

**Gradient boosting on tabular data does not use a GPU.** A `p4d.24xlarge`
would run this model at the same speed for roughly 170 times the price — a
figure the runnable half computes.

The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:

"""Experiments 11 and 14 -- build a model on a managed ML platform, and use
an AutoML service.

SageMaker, Azure ML Studio and Vertex AI need an account, so
`11_sagemaker_train.md` and `14_automl.md` carry the console steps and the
SDK calls, marked as not run here.

What runs here is the model, on the same scikit-learn Course 12 A used --
because the ALGORITHM is not what the cloud changes. What the cloud changes
is the packaging: where the data comes from, where the artefact goes, and
what it costs. Those three are modelled explicitly.

The AutoML half is a genuine (small) AutoML: a real search over real models
with real cross-validation, so the leaderboard is measured. That makes the
point AutoML marketing does not -- what it actually does, and what it costs.
"""
import io
import json
import os
import tempfile
import time

import joblib
import numpy as np
from sklearn.datasets import make_classification
from sklearn.dummy import DummyClassifier
from sklearn.ensemble import GradientBoostingClassifier, RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score, roc_auc_score
from sklearn.model_selection import cross_val_score, train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.tree import DecisionTreeClassifier

import fixtures as f

SEED = 42
MODEL_PATH = os.path.join(tempfile.gettempdir(), "cloud13b_model.joblib")


def churn_data(n=1200):
    """A churn-shaped dataset with a KNOWN base rate, as in Course 12 A."""
    X, y = make_classification(
        n_samples=n, n_features=10, n_informative=5, n_redundant=2,
        weights=[0.85, 0.15], flip_y=0.02, class_sep=1.1,
        random_state=SEED)
    return X, y


def train_model(X_train, y_train):
    """The 'training job'. On SageMaker this is a container; here it is a call."""
    pipe = Pipeline([("scale", StandardScaler()),
                     ("clf", GradientBoostingClassifier(random_state=SEED))])
    pipe.fit(X_train, y_train)
    return pipe


def main():
    print("  Experiments 11 and 14 -- training and AutoML on a managed platform")

    # Step 1: Make the data
    X, y = churn_data()
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.25, stratify=y, random_state=SEED)
    base_rate = y.mean()
    print(f"\n    dataset: {len(X):,} rows, {X.shape[1]} features, "
          f"base rate {base_rate:.2%} positive")
    assert 0.13 < base_rate < 0.17

    # Step 2: Experiment 11: run the training job
    print("\n    --- experiment 11: the training job")
    t0 = time.perf_counter()
    model = train_model(X_train, y_train)
    train_seconds = time.perf_counter() - t0
    pred = model.predict(X_test)
    proba = model.predict_proba(X_test)[:, 1]
    acc = accuracy_score(y_test, pred)
    f1 = f1_score(y_test, pred)
    auc = roc_auc_score(y_test, proba)

    dummy = DummyClassifier(strategy="most_frequent").fit(X_train, y_train)
    dummy_acc = accuracy_score(y_test, dummy.predict(X_test))

    print(f"      {'model':<26}{'accuracy':>10}{'F1':>8}{'AUC':>8}")
    print(f"      {'DummyClassifier':<26}{dummy_acc:>10.4f}"
          f"{0.0:>8.4f}{0.5:>8.4f}")
    print(f"      {'GradientBoosting':<26}{acc:>10.4f}{f1:>8.4f}{auc:>8.4f}")
    assert acc > dummy_acc and auc > 0.85
    print(f"""         quote the DUMMY FIRST, always. {acc:.2%} sounds excellent
         until you see that predicting 'never churns' scores {dummy_acc:.2%}
         -- the real gain is {acc - dummy_acc:.1%} points of accuracy, and it is the
         F1 of {f1:.4f} against 0.0000 that shows the model found anything
         at all. This is Course 12 A's argument, and it does not stop
         being true because the model trained on somebody else's
         computer""")

    # Step 3: Save the artefact, and reload it
    joblib.dump(model, MODEL_PATH)
    size = os.path.getsize(MODEL_PATH)
    reloaded = joblib.load(MODEL_PATH)
    assert (reloaded.predict(X_test) == pred).all()
    print(f"\n      model artefact: {size:,} bytes, reloads and predicts "
          f"identically")
    print("""         THAT FILE IS THE DELIVERABLE. A SageMaker training job
         writes exactly this to s3://bucket/models/, and the deploy
         step reads it back. Training and serving are separate
         systems joined by one artefact in object storage -- which is
         why the IAM role in experiment 10 needs s3:PutObject on
         models/ and nothing else""")

    # Step 4: See what the cloud changes
    print("\n    what a managed platform changes, and what it does not:")
    print(f"      {'':<24}{'your laptop':<24}{'managed platform'}")
    for label, local, cloud in (
            ("the algorithm", "scikit-learn", "SCIKIT-LEARN -- identical"),
            ("data source", "a local file", "s3:// or a feature store"),
            ("hardware", "what you own", "chosen per job, per hour"),
            ("training time", "hours on CPU", "minutes on GPU, if it helps"),
            ("experiment tracking", "a notebook cell", "logged automatically"),
            ("the artefact", "a file you might lose", "versioned in object storage"),
            ("deployment", "you build a server", "one API call"),
            ("cost", "sunk", "PER SECOND, and visible")):
        print(f"      {label:<24}{local:<24}{cloud}")
    print(f"""         THE FIRST ROW IS THE POINT. The cloud does not make your
         model better; it makes training REPRODUCIBLE, deployment
         ROUTINE and cost VISIBLE. A bad model trained on 8 GPUs is
         still a bad model, and the {dummy_acc:.0%} baseline above is unmoved by
         any amount of hardware""")

    # Step 5: Price the instance
    print(f"\n    what this training job would cost (it took "
          f"{train_seconds:.2f}s here):")
    print(f"      {'instance':<16}{'$/hour':>9}{'10 min job':>12}"
          f"{'100 jobs':>11}")
    for inst in ("t3.medium", "m5.xlarge", "c5.4xlarge", "p3.2xlarge",
                 "p4d.24xlarge"):
        rate = f.EC2[inst]
        job = rate / 6
        print(f"      {inst:<16}{rate:>9.4f}{job:>12.4f}{job * 100:>11.2f}")
    cheap = f.EC2["m5.xlarge"] / 6
    gpu = f.EC2["p4d.24xlarge"] / 6
    assert gpu / cheap > 100
    print(f"""         the 8-GPU box costs {gpu / cheap:.0f}x the general-purpose one for the
         same ten minutes. GPUs earn that on deep learning, where the
         work is dense matrix multiplication. GRADIENT BOOSTING ON
         TABULAR DATA DOES NOT USE THEM -- this model would run at the
         same speed and 170x the price.
         'Which instance?' is answered by the ALGORITHM, not by
         ambition""")

    # Step 6: Experiment 14: run AutoML
    print("\n    --- experiment 14: AutoML, actually run")
    candidates = {
        "LogisticRegression": Pipeline([
            ("s", StandardScaler()),
            ("c", LogisticRegression(max_iter=2000, random_state=SEED))]),
        "DecisionTree(d=3)": DecisionTreeClassifier(max_depth=3,
                                                    random_state=SEED),
        "DecisionTree(d=None)": DecisionTreeClassifier(random_state=SEED),
        "RandomForest(100)": RandomForestClassifier(n_estimators=100,
                                                    random_state=SEED),
        "GradientBoosting": GradientBoostingClassifier(random_state=SEED),
    }
    print(f"      {len(candidates)} candidates, 5-fold CV on ROC AUC "
          f"-- a real search")
    board = []
    total_fits = 0
    t0 = time.perf_counter()
    for name, est in candidates.items():
        scores = cross_val_score(est, X_train, y_train, cv=5,
                                 scoring="roc_auc")
        total_fits += 5
        board.append((name, scores.mean(), scores.std()))
    search_seconds = time.perf_counter() - t0
    board.sort(key=lambda r: -r[1])

    print(f"\n      {'rank':<6}{'model':<24}{'CV AUC':>9}{'std':>8}")
    for i, (name, mean, sd) in enumerate(board, 1):
        print(f"      {i:<6}{name:<24}{mean:>9.4f}{sd:>8.4f}")
    winner, best, best_sd = board[0]
    second, second_score, second_sd = board[1]
    print(f"\n      {total_fits} model fits in {search_seconds:.1f}s")
    assert len(board) == 5

    gap = best - second_score
    print(f"""         THE LEADERBOARD IS THE WHOLE OF AUTOML. It fits many
         models, cross-validates each, and ranks them. There is no
         intelligence in it -- it is a SEARCH, and its value is that
         it is exhaustive where you would have been lazy.
         And look at the top two: {best:.4f} against {second_score:.4f}, a gap of
         {gap:.4f} with standard deviations of {best_sd:.4f} and {second_sd:.4f}. THE
         DIFFERENCE IS INSIDE THE NOISE. Declaring a winner here is
         not supported by the data, and 'AutoML picked X' is not a
         reason to prefer X""")

    # Step 7: List what AutoML cannot do
    print("\n    what AutoML does NOT do:")
    for what in (
        "decide what the target variable should be",
        "notice that your target leaks the answer",
        "tell you the base rate matters more than the algorithm",
        "know that last year's data no longer describes this year",
        "choose a threshold that fits the business cost of an error",
        "explain a prediction to a regulator",
        "notice that the model is unfair to a protected group",
    ):
        print(f"      - {what}")
    print("""         EVERY ONE OF THOSE IS THE ACTUAL JOB. AutoML automates
         the part a competent person does in an afternoon and leaves
         untouched the parts that take weeks and cause the failures.
         Say that when asked to evaluate AutoML, and say it before
         saying it is useful -- which it is""")

    # Step 8: Price the search
    print("\n    what an AutoML search costs")
    per_fit_here = search_seconds / total_fits
    print(f"      one fit on THIS dataset (1,200 rows): "
          f"{per_fit_here:.3f} s -- too small to cost anything")
    print("      so scale it: assume one fit takes 4 minutes, which is "
          "ordinary")
    print(f"\n      {'search':<26}{'fits':>6}{'compute':>12}{'m5.xlarge':>12}")
    MINUTES_PER_FIT = 4
    for label, fits in (("this search, scaled", total_fits),
                        ("a modest managed search", 250),
                        ("a full AutoML run", 2000)):
        hours = fits * MINUTES_PER_FIT / 60
        cost = hours * f.EC2["m5.xlarge"]
        print(f"      {label:<26}{fits:>6}{hours:>10.1f} h"
              f"   ${cost:>9.2f}")
    one = 1 * MINUTES_PER_FIT / 60 * f.EC2["m5.xlarge"]
    full = 2000 * MINUTES_PER_FIT / 60 * f.EC2["m5.xlarge"]
    assert full / one == 2000
    print(f"""         AutoML's compute is a straight MULTIPLE of one fit --
         ${full:,.2f} against ${one:.4f}, exactly {full / one:,.0f}x, because that is
         all it is. And AutoML services charge a premium on top of
         the compute.
         So cut the search space with what you already know: on
         tabular data, gradient boosting wins often enough that
         starting there and stopping is frequently the better trade.
         The leaderboard above makes the point -- it spent 5x the
         compute to rank a model {gap:.4f} AUC above the one you would
         have picked anyway, inside the noise""")

    os.remove(MODEL_PATH)
    return board


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 11_sagemaker_train.md:

NOT RUN HERE

11_sagemaker_train.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 11_train_and_automl.py, for experiments 11 and 14:

OUTPUT

  Experiments 11 and 14 -- training and AutoML on a managed platform

    dataset: 1,200 rows, 10 features, base rate 15.67% positive

    --- experiment 11: the training job
      model                       accuracy      F1     AUC
      DummyClassifier               0.8433  0.0000  0.5000
      GradientBoosting              0.9467  0.8095  0.9029
         quote the DUMMY FIRST, always. 94.67% sounds excellent
         until you see that predicting 'never churns' scores 84.33%
         -- the real gain is 10.3% points of accuracy, and it is the
         F1 of 0.8095 against 0.0000 that shows the model found anything
         at all. This is Course 12 A's argument, and it does not stop
         being true because the model trained on somebody else's
         computer

      model artefact: 138,945 bytes, reloads and predicts identically
         THAT FILE IS THE DELIVERABLE. A SageMaker training job
         writes exactly this to s3://bucket/models/, and the deploy
         step reads it back. Training and serving are separate
         systems joined by one artefact in object storage -- which is
         why the IAM role in experiment 10 needs s3:PutObject on
         models/ and nothing else

    what a managed platform changes, and what it does not:
                              your laptop             managed platform
      the algorithm           scikit-learn            SCIKIT-LEARN -- identical
      data source             a local file            s3:// or a feature store
      hardware                what you own            chosen per job, per hour
      training time           hours on CPU            minutes on GPU, if it helps
      experiment tracking     a notebook cell         logged automatically
      the artefact            a file you might lose   versioned in object storage
      deployment              you build a server      one API call
      cost                    sunk                    PER SECOND, and visible
         THE FIRST ROW IS THE POINT. The cloud does not make your
         model better; it makes training REPRODUCIBLE, deployment
         ROUTINE and cost VISIBLE. A bad model trained on 8 GPUs is
         still a bad model, and the 84% baseline above is unmoved by
         any amount of hardware

    what this training job would cost (it took 0.27s here):
      instance           $/hour  10 min job   100 jobs
      t3.medium          0.0416      0.0069       0.69
      m5.xlarge          0.1920      0.0320       3.20
      c5.4xlarge         0.6800      0.1133      11.33
      p3.2xlarge         3.0600      0.5100      51.00
      p4d.24xlarge      32.7726      5.4621     546.21
         the 8-GPU box costs 171x the general-purpose one for the
         same ten minutes. GPUs earn that on deep learning, where the
         work is dense matrix multiplication. GRADIENT BOOSTING ON
         TABULAR DATA DOES NOT USE THEM -- this model would run at the
         same speed and 170x the price.
         'Which instance?' is answered by the ALGORITHM, not by
         ambition

    --- experiment 14: AutoML, actually run
      5 candidates, 5-fold CV on ROC AUC -- a real search

      rank  model                      CV AUC     std
      1     RandomForest(100)          0.9334  0.0210
      2     GradientBoosting           0.9288  0.0196
      3     DecisionTree(d=None)       0.8213  0.0425
      4     LogisticRegression         0.8154  0.0606
      5     DecisionTree(d=3)          0.8025  0.0364

      25 model fits in 2.2s
         THE LEADERBOARD IS THE WHOLE OF AUTOML. It fits many
         models, cross-validates each, and ranks them. There is no
         intelligence in it -- it is a SEARCH, and its value is that
         it is exhaustive where you would have been lazy.
         And look at the top two: 0.9334 against 0.9288, a gap of
         0.0047 with standard deviations of 0.0210 and 0.0196. THE
         DIFFERENCE IS INSIDE THE NOISE. Declaring a winner here is
         not supported by the data, and 'AutoML picked X' is not a
         reason to prefer X

    what AutoML does NOT do:
      - decide what the target variable should be
      - notice that your target leaks the answer
      - tell you the base rate matters more than the algorithm
      - know that last year's data no longer describes this year
      - choose a threshold that fits the business cost of an error
      - explain a prediction to a regulator
      - notice that the model is unfair to a protected group
         EVERY ONE OF THOSE IS THE ACTUAL JOB. AutoML automates
         the part a competent person does in an afternoon and leaves
         untouched the parts that take weeks and cause the failures.
         Say that when asked to evaluate AutoML, and say it before
         saying it is useful -- which it is

    what an AutoML search costs
      one fit on THIS dataset (1,200 rows): 0.089 s -- too small to cost anything
      so scale it: assume one fit takes 4 minutes, which is ordinary

      search                      fits     compute   m5.xlarge
      this search, scaled           25       1.7 h   $     0.32
      a modest managed search      250      16.7 h   $     3.20
      a full AutoML run           2000     133.3 h   $    25.60
         AutoML's compute is a straight MULTIPLE of one fit --
         $25.60 against $0.0128, exactly 2,000x, because that is
         all it is. And AutoML services charge a premium on top of
         the compute.
         So cut the search space with what you already know: on
         tabular data, gradient boosting wins often enough that
         starting there and stopping is frequently the better trade.
         The leaderboard above makes the point -- it spent 5x the
         compute to rank a model 0.0047 AUC above the one you would
         have picked anyway, inside the noise

The artefact is the deliverable. 138,945 bytes, written to disk, reloaded, and predicting identically. A SageMaker training job writes exactly this to s3://bucket/models/, and the deploy step reads it back. Training and serving are separate systems joined by one file in object storage — which is why the IAM role in experiment 10 needs s3:PutObject on models/ and nothing else.

The instance choice, priced:

Instance $/hour 10-min job
m5.xlarge 0.1920 0.0320
p3.2xlarge (1 GPU) 3.0600 0.5100
p4d.24xlarge (8 GPU) 32.7726 5.4621

171× for the same ten minutes — and gradient boosting on tabular data has no GPU code path. It would run at exactly the same speed.

NOTE

"Which instance?" is answered by the algorithm, not by ambition.

The three times the program prints — the training job, the 25 fits, and one fit — measure this machine at one moment and differ from run to run. Corrected: this page said one fit takes 0.127 s; that was one run's figure.

RESULT

Gradient boosting reaches 94.67% against the dummy's 84.33%, F1 0.8095 against 0; the artefact reloads and predicts identically.

Experiment 12 — An ETL job into a cloud warehouse

1. Question

Run a simple ETL job: extract, transform, and load into a cloud warehouse.

2. Aim

Build the job in Glue, load Redshift or BigQuery, and reconcile the counts.

3. Steps

The procedure, on the console and CLI, 12_etl_to_warehouse.md:

  1. Build a Glue job.
  2. Run the crawler.
  3. Load Redshift.
  4. Load BigQuery.
  5. Reconcile the counts.
  6. Orchestrate it.

ETL AGAINST ELT

ELT won because warehouse compute got cheap and elastic. Landing raw data means a transformation bug is fixed by re-running SQL rather than re-extracting from a production database that may no longer hold the old rows — exactly the DELETE problem Big Data Technologies found in Sqoop.

4. Programme

The procedure, on the console and CLI, 12_etl_to_warehouse.md:

# Experiment 12 -- a simple ETL job: extract, transform, load into a cloud warehouse

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`09_etl_warehouse.py`, which runs the whole pipeline into a real columnar warehouse**.

---

<!-- Step 1: Build a Glue job -->
## AWS Glue (the managed option)

```python
import sys
from awsglue.transforms import *
from awsglue.context import GlueContext
from pyspark.context import SparkContext

glueContext = GlueContext(SparkContext.getOrCreate())

src = glueContext.create_dynamic_frame.from_catalog(
    database="retail", table_name="orders")

mapped = ApplyMapping.apply(frame=src, mappings=[
    ("order_id", "long", "order_id", "long"),
    ("region", "string", "region", "string"),
    ("qty", "long", "qty", "int"),
    ("list_price", "double", "revenue_unit", "double"),
])
clean = DropNullFields.apply(frame=mapped)

glueContext.write_dynamic_frame.from_options(
    frame=clean, connection_type="s3",
    connection_options={"path": "s3://retail-lake/curated/",
                        "partitionKeys": ["quarter"]},
    format="parquet")
```

**`partitionKeys` and `format="parquet"` are the two lines that matter**, and
they are Course 12 B's partition pruning and columnar storage arriving in the
cloud unchanged.

<!-- Step 2: Run the crawler -->
## The crawler, and its trap

```bash
aws glue start-crawler --name retail-crawler
```

A crawler infers the schema and registers the table in the Data Catalog.
**It infers types from a sample**, so a column that is integer in the sample
and text later becomes a runtime failure. For anything that matters, define
the table explicitly.

<!-- Step 3: Load Redshift -->
## Loading Redshift

```sql
COPY sales FROM 's3://retail-lake/curated/'
IAM_ROLE 'arn:aws:iam::<acct>:role/RedshiftCopyRole'
FORMAT AS PARQUET;
```

**`COPY` is the only sensible way to load Redshift.** A loop of `INSERT`
statements is orders of magnitude slower, because a columnar store is built
for bulk loads and each `INSERT` writes a nearly-empty block.

```sql
ANALYZE sales;    -- statistics, or the optimiser guesses
VACUUM sales;     -- reclaim space and re-sort after deletes
```

<!-- Step 4: Load BigQuery -->
## Loading BigQuery

```bash
bq load --source_format=PARQUET \
  --time_partitioning_field=order_date \
  --clustering_fields=region,category \
  retail.sales gs://retail-lake/curated/*.parquet
```

**Partition on the column you filter by; cluster on the columns you filter
and group by after that.** Partitioning removes whole days from the scan;
clustering sorts within a partition so the block statistics can skip more.
Both directly reduce **bytes scanned**, which is directly the bill.

<!-- Step 5: Reconcile the counts -->
## The reconciliation nobody does and everybody should

```sql
SELECT COUNT(*) AS rows, SUM(revenue) AS revenue FROM sales;
```

Compare against the source. **Row count and the sum of a money column** —
the same two checks Course 12 B used on the Sqoop import. Rows alone will not
catch a truncated numeric type.

<!-- Step 6: Orchestrate it -->
## Orchestration

| Tool | Shape |
|---|---|
| Step Functions | a state machine, JSON, AWS-native |
| **Airflow (MWAA/Composer)** | a Python DAG — the industry standard |
| Glue Workflows | Glue jobs only |
| EventBridge + Lambda | event-driven, for small steps |

**Retries and idempotency are the whole design problem.** A pipeline that
cannot be re-run safely will eventually be re-run anyway, at 3 a.m., by
someone who does not know what it does.

5. Execution and Results

The procedure, on the console and CLI, 12_etl_to_warehouse.md:

NOT RUN HERE

12_etl_to_warehouse.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model for this experiment, 09_etl_warehouse.py, is shown in full under Experiment 9, with what it printed.

RESULT

The same pipeline as Experiment 9, on managed services; ELT won because warehouse compute got cheap.

Experiment 13 — Monitoring, alarms and auto-scaling

1. Question

Use CloudWatch or Stackdriver to monitor endpoints, set alarms and configure auto-scaling.

2. Aim

Simulate a day of traffic, autoscale it, tune the thresholds, and choose what to alarm on.

3. Steps

The procedure, on the console and CLI, 13_monitoring.md:

  1. Create an alarm.
  2. Set the billing alarm first.
  3. Autoscale an endpoint.
  4. Read what the runnable half shows.
  5. Choose the six metrics.

The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:

  1. Read a day of traffic.
  2. Fix the capacity at the peak.
  3. Autoscale.
  4. Read it honestly.
  5. Tune the thresholds.
  6. Choose what to alarm on.
  7. Price the day.

THE DAY

A day of traffic: peak 1,000 req/s, trough 164; instances serve 150 req/s; the group is 2–12.

Strategy Instance-hours Dropped
fixed at peak (7) 168 0
autoscaled, 70%/40%, cooldown 1 129 1,014

4. Programme

The procedure, on the console and CLI, 13_monitoring.md:

# Experiment 13 -- use CloudWatch/Stackdriver to monitor endpoints, set alarms and auto-scale

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`13_monitoring_autoscale.py`, which runs the control loop and measures what tuning costs**.

---

<!-- Step 1: Create an alarm -->
## An alarm

```bash
aws cloudwatch put-metric-alarm \
  --alarm-name endpoint-p99-latency \
  --namespace AWS/SageMaker \
  --metric-name ModelLatency \
  --dimensions Name=EndpointName,Value=churn-endpoint \
                Name=VariantName,Value=AllTraffic \
  --extended-statistic p99 \
  --period 60 --evaluation-periods 3 \
  --threshold 500000 --comparison-operator GreaterThanThreshold \
  --treat-missing-data notBreaching \
  --alarm-actions arn:aws:sns:us-east-1:<acct>:oncall
```

**`--extended-statistic p99`, not `--statistic Average`.** The runnable half
shows twenty latencies where one request took 900 ms: the mean is 85 ms and
would never fire an alarm, while p99 is 737 ms.

**`ModelLatency` is in MICROSECONDS.** 500000 is half a second. Getting the
unit wrong is how an alarm is set 1,000x too high and never fires.

**`--treat-missing-data`** decides what "no data" means. `notBreaching` is
right for a bursty endpoint; `breaching` is right when silence itself is the
failure.

<!-- Step 2: Set the billing alarm first -->
## The billing alarm, which comes first

```bash
aws cloudwatch put-metric-alarm --alarm-name monthly-spend \
  --namespace AWS/Billing --metric-name EstimatedCharges \
  --dimensions Name=Currency,Value=USD \
  --statistic Maximum --period 21600 --evaluation-periods 1 \
  --threshold 50 --comparison-operator GreaterThanThreshold \
  --alarm-actions arn:aws:sns:us-east-1:<acct>:oncall
```

**Billing metrics exist only in `us-east-1`** regardless of where you work,
and they lag by about six hours. Set this before anything else in the course.

<!-- Step 3: Autoscale an endpoint -->
## Auto-scaling a SageMaker endpoint

```bash
aws application-autoscaling register-scalable-target \
  --service-namespace sagemaker \
  --resource-id endpoint/churn-endpoint/variant/AllTraffic \
  --scalable-dimension sagemaker:variant:DesiredInstanceCount \
  --min-capacity 1 --max-capacity 8

aws application-autoscaling put-scaling-policy \
  --policy-name track-invocations --policy-type TargetTrackingScaling \
  --service-namespace sagemaker \
  --resource-id endpoint/churn-endpoint/variant/AllTraffic \
  --scalable-dimension sagemaker:variant:DesiredInstanceCount \
  --target-tracking-scaling-policy-configuration '{
     "TargetValue": 750.0,
     "PredefinedMetricSpecification":
        {"PredefinedMetricType": "SageMakerVariantInvocationsPerInstance"},
     "ScaleInCooldown": 300, "ScaleOutCooldown": 60}'
```

**`ScaleOutCooldown` short and `ScaleInCooldown` long.** Scale out eagerly
because being under capacity drops requests; scale in reluctantly because
scaling back out costs boot time. The runnable half measures what happens
when you get that backwards.

<!-- Step 4: Read what the runnable half shows -->
## What the runnable half shows, and it is not flattering

- **Autoscaling dropped 1,014 requests where fixed capacity dropped none**,
  because the group is always sized for the previous observation.
- **The most aggressive configuration cost MORE than fixed capacity** —
  188 instance-hours against 168 — by overshooting.

**"Autoscaling saves money" is a claim about a tuned autoscaler.** Say that,
and give the numbers.

<!-- Step 5: Choose the six metrics -->
## The six metrics worth alarming on

| Metric | Alarm when | The trap |
|---|---|---|
| `ModelLatency` p99 | > 500 ms for 3 min | the mean hides it |
| `Invocation5XXErrors` | > 0 for 1 min | these are **yours** |
| `Invocation4XXErrors` | > 1% of requests | a rate, never a count |
| `CPUUtilization` | > 70% for 5 min | I/O-bound apps never reach it |
| `EstimatedCharges` | > your budget | lags ~6 h, `us-east-1` only |
| **`Invocations` == 0** | for 1 hour | **a dead endpoint still bills** |

**The last row is the one people miss.** An endpoint serving nothing looks
perfect on every performance metric and costs exactly the same as a busy one.

The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:

"""Experiment 13 -- CloudWatch/Stackdriver: monitor an endpoint, set alarms,
and configure auto-scaling rules.

The console steps are in `13_monitoring.md`, marked as not run here.

What runs here is the CONTROL LOOP, which is the part that actually behaves
in ways people do not expect: scaling lags demand, aggressive thresholds
oscillate, cooldowns trade responsiveness for stability, and an alarm on an
average hides the tail. Every one of those is measured below rather than
asserted.
"""
import fixtures as f

CAPACITY_PER_INSTANCE = 150      # requests/sec one instance can serve
MIN_INSTANCES = 2
MAX_INSTANCES = 12


def simulate(traffic, scale_out_at, scale_in_at, cooldown,
             start=MIN_INSTANCES, step=1):
    """A target-tracking autoscaler, one tick per hour.

    The key realism: a scaling decision made at tick t only takes effect at
    tick t+1. Real instances take minutes to boot, so the group is ALWAYS
    sized for the PREVIOUS observation.
    """
    instances = start
    cool = 0
    history = []
    for demand in traffic:
        capacity = instances * CAPACITY_PER_INSTANCE
        util = demand / capacity
        served = min(demand, capacity)
        dropped = demand - served
        history.append({"demand": demand, "instances": instances,
                        "util": util, "dropped": dropped})
        if cool > 0:
            cool -= 1
        elif util > scale_out_at and instances < MAX_INSTANCES:
            instances = min(MAX_INSTANCES, instances + step)
            cool = cooldown
        elif util < scale_in_at and instances > MIN_INSTANCES:
            instances = max(MIN_INSTANCES, instances - step)
            cool = cooldown
    return history


def summarise(history):
    hours = len(history)
    return {
        "instance_hours": sum(h["instances"] for h in history),
        "dropped": sum(h["dropped"] for h in history),
        "peak_instances": max(h["instances"] for h in history),
        "hours_over_90": sum(1 for h in history if h["util"] > 0.90),
        "mean_util": sum(h["util"] for h in history) / hours,
        "changes": sum(1 for a, b in zip(history, history[1:])
                       if a["instances"] != b["instances"]),
    }


def percentile(values, p):
    s = sorted(values)
    k = (len(s) - 1) * p / 100
    lo, hi = int(k), min(int(k) + 1, len(s) - 1)
    return s[lo] + (s[hi] - s[lo]) * (k - lo)


def main():
    print("  Experiment 13 -- monitoring, alarms and auto-scaling")

    # Step 1: Read a day of traffic
    traffic = f.daily_traffic()
    print(f"\n    a day of traffic: peak {max(traffic)} req/s, "
          f"trough {min(traffic)} req/s, {sum(traffic):,} req/s-hours")
    print(f"    one instance serves {CAPACITY_PER_INSTANCE} req/s; "
          f"group is {MIN_INSTANCES}..{MAX_INSTANCES}")

    # Step 2: Fix the capacity at the peak
    print("\n    OPTION 1 -- fixed capacity, sized for peak:")
    need = -(-max(traffic) // CAPACITY_PER_INSTANCE)
    fixed = [{"demand": d, "instances": need,
              "util": d / (need * CAPACITY_PER_INSTANCE),
              "dropped": 0} for d in traffic]
    fs = summarise(fixed)
    print(f"      {need} instances all day = {fs['instance_hours']} "
          f"instance-hours, 0 dropped")
    print(f"      mean utilisation {fs['mean_util'] * 100:.1f}%")
    assert fs["dropped"] == 0
    print("""         nothing is dropped and most of the fleet is idle most of
         the day. That is the pre-cloud bargain: you buy the peak and
         pay for it at 3 a.m.""")

    # Step 3: Autoscale
    print("\n    OPTION 2 -- target tracking, scale out above 70%, "
          "in below 40%, cooldown 1:")
    auto = simulate(traffic, 0.70, 0.40, cooldown=1)
    a = summarise(auto)
    print(f"      {'hour':>5}{'demand':>8}{'inst':>6}{'util':>8}{'dropped':>9}")
    for h, row in enumerate(auto):
        flag = "  <-- shortfall" if row["dropped"] else ""
        print(f"      {h:>5}{row['demand']:>8}{row['instances']:>6}"
              f"{row['util'] * 100:>7.0f}%{row['dropped']:>9}{flag}")
    saved = fs["instance_hours"] - a["instance_hours"]
    pct = 100 * saved / fs["instance_hours"]
    print(f"\n      instance-hours {a['instance_hours']} against "
          f"{fs['instance_hours']} fixed -- {pct:.0f}% fewer")
    print(f"      requests dropped: {a['dropped']:,}")
    print(f"      scaling changes : {a['changes']}")
    assert a["instance_hours"] < fs["instance_hours"]
    assert a["dropped"] > 0

    # Step 4: Read it honestly
    worst = max(auto, key=lambda r: r["dropped"])
    hour = auto.index(worst)
    print(f"""
         AUTOSCALING DROPPED {a['dropped']:,} REQUESTS AND FIXED CAPACITY DROPPED
         NONE. The worst hour is {hour}, where demand jumped to {worst['demand']}
         against {worst['instances']} instances -- the group was sized for the
         PREVIOUS hour, because a scaling decision takes effect one
         tick late.
         Autoscaling does not track demand. It CHASES demand, and it
         is always one observation behind. That lag is the cost of the
         {pct:.0f}% saving, and pretending otherwise is how a launch goes
         badly""")

    # Step 5: Tune the thresholds
    print("\n    the same day at different thresholds:")
    print(f"      {'out/in':<12}{'cool':>5}{'inst-hrs':>10}{'dropped':>9}"
          f"{'changes':>9}{'mean util':>11}")
    configs = [(0.70, 0.40, 1), (0.50, 0.30, 1), (0.85, 0.60, 1),
               (0.70, 0.40, 3), (0.50, 0.30, 0)]
    results = {}
    for out, inn, cd in configs:
        r = summarise(simulate(traffic, out, inn, cd))
        results[(out, inn, cd)] = r
        print(f"      {f'{out:.0%}/{inn:.0%}':<12}{cd:>5}"
              f"{r['instance_hours']:>10}{r['dropped']:>9}"
              f"{r['changes']:>9}{r['mean_util'] * 100:>10.0f}%")

    aggressive = results[(0.50, 0.30, 1)]
    lazy = results[(0.85, 0.60, 1)]
    assert aggressive["dropped"] < lazy["dropped"]
    assert aggressive["instance_hours"] > lazy["instance_hours"]
    print(f"""         THE TABLE IS A TRADE-OFF CURVE, NOT A LEADERBOARD.
         Scaling out at 50% drops {aggressive['dropped']:,} requests for {aggressive['instance_hours']} instance-hours;
         scaling out at 85% drops {lazy['dropped']:,} for {lazy['instance_hours']}. You are choosing
         between spare capacity and dropped requests, and only a
         business can say which is worse""")

    no_cool = results[(0.50, 0.30, 0)]
    assert no_cool["dropped"] == 0
    assert no_cool["instance_hours"] > fs["instance_hours"]
    print(f"""
         AND READ THE LAST ROW AGAINST FIXED CAPACITY. Scaling out at
         50% with NO cooldown drops nothing -- and costs {no_cool['instance_hours']}
         instance-hours against fixed capacity's {fs['instance_hours']}.
         AUTOSCALING MADE IT MORE EXPENSIVE. Chase demand hard enough
         and the group overshoots on the way up and lingers on the way
         down, so you buy more than the peak. 'Autoscaling saves
         money' is a claim about a TUNED autoscaler, not about
         autoscaling""")

    long_cool = results[(0.70, 0.40, 3)]
    short_cool = results[(0.70, 0.40, 1)]
    print(f"""
         and the cooldown: 3 ticks gives {long_cool['changes']} scaling changes against
         {short_cool['changes']}, at {long_cool['dropped'] - short_cool['dropped']:+,} dropped requests. A long cooldown
         stops FLAPPING -- scaling out and back in repeatedly around a
         threshold, which costs boot time and stabilises nothing""")

    # Step 6: Choose what to alarm on
    print("\n    alarms: what to measure, and the trap in each")
    lat = [40, 42, 41, 45, 43, 40, 44, 42, 41, 43,
           41, 42, 40, 44, 43, 41, 42, 45, 40, 900]
    mean = sum(lat) / len(lat)
    p50, p95, p99 = (percentile(lat, p) for p in (50, 95, 99))
    print(f"      20 request latencies, one of them 900 ms:")
    print(f"        mean {mean:.1f} ms   p50 {p50:.1f} ms   "
          f"p95 {p95:.1f} ms   p99 {p99:.1f} ms")
    assert mean < 100 and p99 > 400
    print(f"""         AN ALARM ON THE MEAN ({mean:.0f} ms) NEVER FIRES. One request
         in twenty took 900 ms and the average absorbed it. Alarm on
         p95 or p99, because the tail is where users live -- and note
         that 5% of requests is a lot of users""")

    print(f"\n      {'metric':<22}{'alarm when':<26}{'the trap'}")
    for m, when, trap in (
            ("CPUUtilization", "> 70% for 5 min", "an I/O-bound app never hits it"),
            ("ModelLatency p99", "> 500 ms for 3 min", "the mean hides it"),
            ("Invocation4XXErrors", "> 1% of requests", "a rate, never a raw count"),
            ("Invocation5XXErrors", "> 0 for 1 min", "these are YOUR fault"),
            ("EstimatedCharges", "> your monthly budget", "billing metrics lag ~6 h"),
            ("no invocations at all", "== 0 for 1 hour", "a dead endpoint still bills")):
        print(f"      {m:<22}{when:<26}{trap}")
    print("""         THE LAST ROW IS THE ONE PEOPLE MISS. An endpoint serving
         nothing looks perfect on every performance metric and costs
         the same as a busy one. Alarm on ABSENCE of traffic, and
         alarm on spend -- those two catch the failures that
         monitoring dashboards are blind to""")

    # Step 7: Price the day
    print("\n    what the day cost, at m5.large on-demand:")
    rate = f.EC2["m5.large"]
    for label, hours in (("fixed at peak", fs["instance_hours"]),
                         ("autoscaled", a["instance_hours"])):
        print(f"      {label:<18}{hours:>5} instance-hours  "
              f"${hours * rate:>7.2f}/day   ${hours * rate * 30:>8.2f}/month")
    fixed_month = fs["instance_hours"] * rate * 30
    auto_month = a["instance_hours"] * rate * 30
    spot_month = auto_month * (1 - f.SPOT_DISCOUNT)
    res_month = fixed_month * (1 - f.RESERVED_DISCOUNT)
    print(f"\n      autoscaled on SPOT (-{f.SPOT_DISCOUNT:.0%}, interruptible): "
          f"${spot_month:,.2f}/month")
    print(f"      fixed on RESERVED (-{f.RESERVED_DISCOUNT:.0%}, 1-yr commit): "
          f"${res_month:,.2f}/month")
    assert spot_month < auto_month < fixed_month
    print(f"""         the real answer is usually BOTH: a reserved baseline for
         the floor you always need, autoscaled on-demand or spot for
         the peak. Here that is {MIN_INSTANCES} reserved instances plus the rest
         elastic -- and it beats either pure strategy, which is why
         every cost-optimisation review starts by asking what your
         floor is""")


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 13_monitoring.md:

NOT RUN HERE

13_monitoring.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 13_monitoring_autoscale.py, for experiment 13:

OUTPUT

  Experiment 13 -- monitoring, alarms and auto-scaling

    a day of traffic: peak 1000 req/s, trough 164 req/s, 13,996 req/s-hours
    one instance serves 150 req/s; group is 2..12

    OPTION 1 -- fixed capacity, sized for peak:
      7 instances all day = 168 instance-hours, 0 dropped
      mean utilisation 55.5%
         nothing is dropped and most of the fleet is idle most of
         the day. That is the pre-cloud bargain: you buy the peak and
         pay for it at 3 a.m.

    OPTION 2 -- target tracking, scale out above 70%, in below 40%, cooldown 1:
       hour  demand  inst    util  dropped
          0     208     2     69%        0
          1     182     2     61%        0
          2     164     2     55%        0
          3     164     2     55%        0
          4     173     2     58%        0
          5     226     2     75%        0
          6     366     3     81%        0
          7     604     3    134%      154  <-- shortfall
          8     824     4    137%      224  <-- shortfall
          9    1000     4    167%      400  <-- shortfall
         10     930     5    124%      180  <-- shortfall
         11     806     5    107%       56  <-- shortfall
         12     736     6     82%        0
         13     701     6     78%        0
         14     666     7     63%        0
         15     692     7     66%        0
         16     771     7     73%        0
         17     894     8     74%        0
         18     965     8     80%        0
         19     912     9     68%        0
         20     754     9     56%        0
         21     578     9     43%        0
         22     402     9     30%        0
         23     278     8     23%        0

      instance-hours 129 against 168 fixed -- 23% fewer
      requests dropped: 1,014
      scaling changes : 8

         AUTOSCALING DROPPED 1,014 REQUESTS AND FIXED CAPACITY DROPPED
         NONE. The worst hour is 9, where demand jumped to 1000
         against 4 instances -- the group was sized for the
         PREVIOUS hour, because a scaling decision takes effect one
         tick late.
         Autoscaling does not track demand. It CHASES demand, and it
         is always one observation behind. That lag is the cost of the
         23% saving, and pretending otherwise is how a launch goes
         badly

    the same day at different thresholds:
      out/in       cool  inst-hrs  dropped  changes  mean util
      70%/40%         1       129     1014        8        77%
      50%/30%         1       158      358       10        62%
      85%/60%         1       114     1380        7        87%
      70%/40%         3        96     2093        4        98%
      50%/30%         0       188        0       11        52%
         THE TABLE IS A TRADE-OFF CURVE, NOT A LEADERBOARD.
         Scaling out at 50% drops 358 requests for 158 instance-hours;
         scaling out at 85% drops 1,380 for 114. You are choosing
         between spare capacity and dropped requests, and only a
         business can say which is worse

         AND READ THE LAST ROW AGAINST FIXED CAPACITY. Scaling out at
         50% with NO cooldown drops nothing -- and costs 188
         instance-hours against fixed capacity's 168.
         AUTOSCALING MADE IT MORE EXPENSIVE. Chase demand hard enough
         and the group overshoots on the way up and lingers on the way
         down, so you buy more than the peak. 'Autoscaling saves
         money' is a claim about a TUNED autoscaler, not about
         autoscaling

         and the cooldown: 3 ticks gives 4 scaling changes against
         8, at +1,079 dropped requests. A long cooldown
         stops FLAPPING -- scaling out and back in repeatedly around a
         threshold, which costs boot time and stabilises nothing

    alarms: what to measure, and the trap in each
      20 request latencies, one of them 900 ms:
        mean 85.0 ms   p50 42.0 ms   p95 87.8 ms   p99 737.5 ms
         AN ALARM ON THE MEAN (85 ms) NEVER FIRES. One request
         in twenty took 900 ms and the average absorbed it. Alarm on
         p95 or p99, because the tail is where users live -- and note
         that 5% of requests is a lot of users

      metric                alarm when                the trap
      CPUUtilization        > 70% for 5 min           an I/O-bound app never hits it
      ModelLatency p99      > 500 ms for 3 min        the mean hides it
      Invocation4XXErrors   > 1% of requests          a rate, never a raw count
      Invocation5XXErrors   > 0 for 1 min             these are YOUR fault
      EstimatedCharges      > your monthly budget     billing metrics lag ~6 h
      no invocations at all == 0 for 1 hour           a dead endpoint still bills
         THE LAST ROW IS THE ONE PEOPLE MISS. An endpoint serving
         nothing looks perfect on every performance metric and costs
         the same as a busy one. Alarm on ABSENCE of traffic, and
         alarm on spend -- those two catch the failures that
         monitoring dashboards are blind to

    what the day cost, at m5.large on-demand:
      fixed at peak       168 instance-hours  $  16.13/day   $  483.84/month
      autoscaled          129 instance-hours  $  12.38/day   $  371.52/month

      autoscaled on SPOT (-70%, interruptible): $111.46/month
      fixed on RESERVED (-40%, 1-yr commit): $290.30/month
         the real answer is usually BOTH: a reserved baseline for
         the floor you always need, autoscaled on-demand or spot for
         the peak. Here that is 2 reserved instances plus the rest
         elastic -- and it beats either pure strategy, which is why
         every cost-optimisation review starts by asking what your
         floor is

AUTOSCALING DROPPED 1,014 REQUESTS AND FIXED CAPACITY DROPPED NONE

The worst hour is hour 9: demand jumped to 1,000 against 4 instances, because the group was sized for the previous hour.

NOTE

Autoscaling does not track demand. It CHASES demand, and it is always one observation behind.

That lag is the cost of the 23% saving.

The tuning curve:

Out/in Cool Inst-hrs Dropped Changes
70%/40% 1 129 1,014 8
50%/30% 1 158 358 10
85%/60% 1 114 1,380 7
70%/40% 3 96 2,093 4
50%/30% 0 188 0 11

This is a trade-off curve, not a leaderboard. You are choosing between spare capacity and dropped requests, and only a business can say which is worse.

AND READ THE LAST ROW AGAINST FIXED CAPACITY

188 instance-hours against 168. Scaling out at 50% with no cooldown drops nothing — and costs more than simply buying the peak.

NOTE

Autoscaling made it MORE expensive. Chase demand hard enough and the group overshoots on the way up and lingers on the way down.

"Autoscaling saves money" is a claim about a tuned autoscaler.

The cooldown row: 3 ticks gives 4 scaling changes instead of 8, at +1,079 dropped requests. A long cooldown stops flapping — which costs boot time and stabilises nothing. Scale out eagerly, scale in reluctantly.

Alarm on the tail:

20 latencies, one of them 900 ms:
  mean 85.0 ms   p50 42.0 ms   p95 87.8 ms   p99 737.5 ms

An alarm on the mean never fires. Alarm on p95 or p99 — the tail is where users live, and 5% of requests is a lot of users.

THE METRIC NOBODY SETS

Metric Alarm when The trap
ModelLatency p99 > 500 ms the mean hides it; the unit is microseconds
Invocation5XXErrors > 0 these are yours
Invocation4XXErrors > 1% a rate, never a count
EstimatedCharges > budget lags ~6 h, us-east-1 only
Invocations == 0 for 1 hour a dead endpoint still bills

The last row is the one people miss. An endpoint serving nothing looks perfect on every performance metric and costs the same as a busy one.

And what the day cost:

Per month
fixed at peak, on-demand $483.84
autoscaled, on-demand $371.52
fixed, reserved (−40%) $290.30
autoscaled, spot (−70%) $111.46

The real answer is usually both: a reserved baseline for the floor, spot or on-demand for the peak.

RESULT

Autoscaling saved 23% of instance-hours and dropped 1,014 requests that fixed capacity kept; tuned hard enough, it cost more than fixed capacity.

Experiment 14 — AutoML

1. Question

Use a cloud AutoML service for a prediction task.

2. Aim

Run a model search, read its leaderboard, and see what it does not do.

3. Steps

The procedure, on the console and CLI, 14_automl.md:

  1. Run SageMaker Autopilot.
  2. Or Vertex AI AutoML.
  3. Or Azure Automated ML.
  4. See what they do.
  5. And what they do not.
  6. Read the explainability report.

AUTOML, ACTUALLY RUN

Five candidates, 5-fold CV, 25 real fits:

Rank Model CV AUC std
1 RandomForest(100) 0.9334 0.0210
2 GradientBoosting 0.9288 0.0196
3 DecisionTree(depth=None) 0.8213 0.0425
4 LogisticRegression 0.8154 0.0606
5 DecisionTree(depth=3) 0.8025 0.0364

The leaderboard is the whole of AutoML. Fit many models, cross-validate, rank. There is no intelligence in it — it is a search.

4. Programme

The procedure, on the console and CLI, 14_automl.md:

# Experiment 14 -- use cloud AutoML services for a dataset prediction task

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`11_train_and_automl.py`, which runs a real 5-model, 5-fold search and reports the leaderboard**.

---

<!-- Step 1: Run SageMaker Autopilot -->
## SageMaker Autopilot

```python
from sagemaker.automl.automl import AutoML

automl = AutoML(
    role=role,
    target_attribute_name="churned",
    output_path=f"s3://{bucket}/autopilot/",
    problem_type="BinaryClassification",
    job_objective={"MetricName": "AUC"},
    max_candidates=20,
    max_runtime_per_training_job_in_seconds=600,
)
automl.fit(inputs=f"s3://{bucket}/train/train.csv")
automl.describe_auto_ml_job()["BestCandidate"]
```

**`max_candidates` and `max_runtime_per_training_job_in_seconds` are the
budget**, and they are not optional. Without them the job explores until it
is satisfied, and it bills the whole time.

<!-- Step 2: Or Vertex AI AutoML -->
## Vertex AI AutoML

```bash
gcloud ai custom-jobs create --region=us-central1 ...
# or, tabular:
gcloud beta ai models list --region=us-central1
```

Vertex bills AutoML in **node-hours** with a documented minimum. Read the
minimum before starting — a small dataset does not produce a small bill.

<!-- Step 3: Or Azure Automated ML -->
## Azure Automated ML

```python
from azure.ai.ml import automl
job = automl.classification(
    training_data=train, target_column_name="churned",
    primary_metric="AUC_weighted",
    experiment_timeout_minutes=30,          # THE BUDGET
    enable_early_termination=True,
)
```

<!-- Step 4: See what they do -->
## What these actually do

**They fit many models, cross-validate each, and rank them.** The runnable
half does exactly this with five candidates and 5-fold CV, and prints the
leaderboard. There is no intelligence in it — it is a **search**, and its
value is that it is exhaustive where a human would be lazy.

**And read the top of that leaderboard carefully.** In the run here the top
two models differ by 0.0047 AUC with standard deviations of 0.0210 and
0.0196. **The difference is inside the noise**, and "AutoML picked X" is not
a reason to prefer X.

<!-- Step 5: And what they do not -->
## What AutoML does not do

- decide what the target variable should be
- notice that your target **leaks** the answer
- tell you the base rate matters more than the algorithm
- know that last year's data no longer describes this year
- choose a threshold that fits the business cost of an error
- explain a prediction to a regulator
- notice the model is unfair to a protected group

**Every one of those is the actual job.** AutoML automates the afternoon and
leaves the weeks untouched.

<!-- Step 6: Read the explainability report -->
## The explainability report

Autopilot generates a candidate-definition notebook and a data-exploration
notebook. **Read them** — they are the best thing about the product, because
they show you the feature engineering it chose, which is the part you would
otherwise never see and could not defend.

5. Execution and Results

The procedure, on the console and CLI, 14_automl.md:

NOT RUN HERE

14_automl.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

AND THE TOP TWO ARE INSIDE THE NOISE

0.9334 against 0.9288 is a gap of 0.0047, with standard deviations of 0.0210 and 0.0196. Declaring a winner is not supported by the data, and "AutoML picked X" is not a reason to prefer X.

What the search costs. One fit here takes a fraction of a second — too small to cost anything. Scale to a realistic four minutes per fit:

Search Fits Compute m5.xlarge
this search 25 1.7 h $0.32
a modest managed search 250 16.7 h $3.20
a full AutoML run 2,000 133.3 h $25.60

A straight multiple of one fit — exactly 2,000× — because that is all it is. And managed services charge a premium on top.

WHAT AUTOML DOES NOT DO

decide the target · notice leakage · tell you the base rate matters more · know last year's data no longer applies · choose a threshold that fits the business cost · explain a prediction · notice unfairness

Every one of those is the actual job. AutoML automates the afternoon and leaves the weeks untouched.

The Python model for this experiment, 11_train_and_automl.py, is shown in full under Experiment 11, with what it printed.

RESULT

Five candidates, 25 real fits: random forest leads gradient boosting by 0.0047, inside the noise.

Experiment 15 — Deploy the model as a REST endpoint

1. Question

Deploy a trained ML model as a REST API endpoint.

2. Aim

Serve the model over HTTP, check health, predictions and errors, and measure latency and batching.

3. Steps

The procedure, on the console and CLI, 15_deploy.md:

  1. Deploy.
  2. Invoke it.
  3. Follow the container contract.
  4. Choose real-time, serverless or batch.
  5. Delete it.
  6. Monitor it.

The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:

  1. Train the model, and save it.
  2. Serve it.
  3. Check its health.
  4. Ask for a prediction.
  5. Send a batch.
  6. Send bad requests.
  7. Measure the latency.
  8. Batch, against one at a time.
  9. Read the metrics.
  10. Compare with a managed endpoint.

THE EQUALITY THAT IS THE DEPLOYMENT TEST

artefact loaded from disk: 138,945 bytes

GET /ping        -> 200 {'status': 'healthy', 'model_loaded': True}
POST /invocations (3 rows)  -> 200
  predictions   : [0, 0, 0]
  probabilities : [0.015383, 0.008643, 0.014492]
POST /invocations (300 rows) -> 200, accuracy 0.9467

The endpoint's answers are identical to calling the model in-process. Serving must not change predictions — and a preprocessing step that lives in your notebook rather than in the pipeline is exactly how it does.

4. Programme

The procedure, on the console and CLI, 15_deploy.md:

# Experiment 15 -- deploy a trained ML model as a REST API endpoint

## NOT EXECUTED

**This repository has no cloud account, and none will be created.** Creating
one requires a payment card and accepts a billing relationship, which is not
something a study repository should do on anyone's behalf.

So this file is what you type, with the traps marked. **It has never been
run here**, and nothing in the notes claims an output for it.

The runnable half is **`15_deploy_endpoint.py`, which starts a REAL HTTP server, serves a REAL model and calls it**.

---

<!-- Step 1: Deploy -->
## Deploy

```python
predictor = estimator.deploy(
    initial_instance_count=1,
    instance_type="ml.m5.large",
    endpoint_name="churn-endpoint",
)
predictor.predict([[1.2, 0.4, ...]])
```

Or from the CLI, in the three steps the SDK hides:

```bash
aws sagemaker create-model --model-name churn-model \
  --primary-container Image=<ecr-uri>,ModelDataUrl=s3://bucket/models/model.tar.gz \
  --execution-role-arn arn:aws:iam::<acct>:role/SageMakerExecutionRole

aws sagemaker create-endpoint-config --endpoint-config-name churn-config \
  --production-variants VariantName=AllTraffic,ModelName=churn-model,\
InitialInstanceCount=1,InstanceType=ml.m5.large,InitialVariantWeight=1

aws sagemaker create-endpoint --endpoint-name churn-endpoint \
  --endpoint-config-name churn-config
```

**Model, endpoint config, endpoint — three objects, not one.** That is what
makes blue/green possible: create a second config and update the endpoint,
and traffic shifts without downtime.

<!-- Step 2: Invoke it -->
## Invoke it

```bash
aws sagemaker-runtime invoke-endpoint \
  --endpoint-name churn-endpoint \
  --content-type application/json \
  --body '{"instances": [[1.2, 0.4, 3.3, ...]]}' \
  /dev/stdout
```

**The request is IAM-signed.** There is no API key to leak, and access is
governed by the same policy evaluation as everything else.

<!-- Step 3: Follow the container contract -->
## The container contract

Your container must answer two routes:

| Route | Must |
|---|---|
| `GET /ping` | return 200 quickly, **without running the model** |
| `POST /invocations` | run inference |

**A health check that does real inference marks the container unhealthy
whenever the model is merely slow — and the platform then kills a container
that was working.** The runnable half implements both routes and says so.

<!-- Step 4: Choose real-time, serverless or batch -->
## Real-time, serverless or batch

| | Real-time | Serverless | Batch transform |
|---|---|---|---|
| Latency | ms | ms, **after a cold start** | minutes to hours |
| Billed | **per hour, always** | per request | per job |
| Idle cost | **the full instance** | **zero** | zero |
| Good for | steady traffic | spiky or occasional | scoring a whole file |

**An `ml.m5.large` endpoint is about $70/month whether or not anything calls
it.** If traffic is occasional, serverless inference costs a fraction; if you
are scoring a file, batch transform is the right tool and an endpoint is the
expensive way to do arithmetic — the runnable half measures a **37x**
difference between one batched request and 100 single ones.

<!-- Step 5: Delete it -->
## Then delete it

```bash
aws sagemaker delete-endpoint --endpoint-name churn-endpoint
aws sagemaker delete-endpoint-config --endpoint-config-name churn-config
aws sagemaker delete-model --model-name churn-model
```

**Deleting the endpoint is a step in the experiment, not an afterthought.**
Every "surprise AWS bill" story is a resource nobody switched off.

<!-- Step 6: Monitor it -->
## Monitoring the deployed model

Data drift is the failure that has no error message: the endpoint keeps
returning 200 and the predictions quietly stop being right.

```python
from sagemaker.model_monitor import DefaultModelMonitor
monitor = DefaultModelMonitor(role=role, instance_type="ml.m5.xlarge")
monitor.suggest_baseline(baseline_dataset=f"s3://{bucket}/train/train.csv")
monitor.create_monitoring_schedule(endpoint_input=predictor.endpoint_name,
                                   schedule_cron_expression="cron(0 * ? * * *)")
```

**Baseline the training distribution, then compare production inputs against
it hourly.** A drift alarm is the only thing that catches a model that has
stopped working while every infrastructure metric stays green.

The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:

"""Experiment 15 -- deploy a trained ML model as a REST API endpoint.

`15_deploy.md` carries the SageMaker deploy call and the console steps, NOT
EXECUTED -- there is no cloud account.

But the ENDPOINT ITSELF RUNS HERE. A real HTTP server starts on localhost,
serves a real scikit-learn model over a real JSON API, is called with real
requests, and is shut down. The contract, the error handling, the health
check and the latency measurements are genuine -- only the hosting is not.

That is the honest split: SageMaker gives you a container, a load balancer,
autoscaling and an IAM-signed URL. What it serves is this.
"""
import json
import os
import tempfile
import threading
import time
import urllib.error
import urllib.request
from http.server import BaseHTTPRequestHandler, HTTPServer

import joblib
import numpy as np

import fixtures as f
from sklearn.datasets import make_classification
from sklearn.ensemble import GradientBoostingClassifier
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler

SEED = 42
N_FEATURES = 10
HOST = "127.0.0.1"
PORT = 0          # the OS picks a free port; the real one is read back

_MODEL = None
_STATE = {"invocations": 0, "errors_4xx": 0, "errors_5xx": 0,
          "latencies": []}


class Handler(BaseHTTPRequestHandler):
    """The two routes every model endpoint must have, and no more."""

    def log_message(self, *args):
        pass                                  # keep the suite output clean

    def _send(self, code, payload):
        body = json.dumps(payload).encode()
        self.send_response(code)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(body)))
        self.end_headers()
        self.wfile.write(body)

    def do_GET(self):
        if self.path == "/ping":
            # SageMaker calls THIS to decide whether the container is alive.
            # It must not run the model: a health check that does real work
            # takes the endpoint down when the model is merely slow.
            self._send(200, {"status": "healthy",
                             "model_loaded": _MODEL is not None})
        elif self.path == "/metrics":
            self._send(200, dict(_STATE, latencies=len(_STATE["latencies"])))
        else:
            _STATE["errors_4xx"] += 1
            self._send(404, {"error": "not found"})

    def do_POST(self):
        if self.path != "/invocations":
            _STATE["errors_4xx"] += 1
            self._send(404, {"error": "not found"})
            return
        started = time.perf_counter()
        try:
            length = int(self.headers.get("Content-Length", 0))
            payload = json.loads(self.rfile.read(length) or b"{}")
            rows = payload.get("instances")
            if not isinstance(rows, list) or not rows:
                raise ValueError("body must be {'instances': [[...], ...]}")
            arr = np.asarray(rows, dtype=float)
            if arr.ndim != 2 or arr.shape[1] != N_FEATURES:
                raise ValueError(
                    f"each instance needs {N_FEATURES} features, "
                    f"got shape {list(arr.shape)}")
        except (ValueError, TypeError, json.JSONDecodeError) as exc:
            # A BAD REQUEST IS A 4XX, NOT A 5XX. Getting this wrong makes
            # your error alarm fire for other people's mistakes.
            _STATE["errors_4xx"] += 1
            self._send(400, {"error": str(exc)})
            return
        try:
            proba = _MODEL.predict_proba(arr)[:, 1]
            preds = (proba >= 0.5).astype(int)
        except Exception as exc:                       # pragma: no cover
            _STATE["errors_5xx"] += 1
            self._send(500, {"error": "inference failed"})
            return
        _STATE["invocations"] += len(rows)
        _STATE["latencies"].append((time.perf_counter() - started) * 1000)
        self._send(200, {"predictions": preds.tolist(),
                         "probabilities": [round(p, 6) for p in proba]})


def build_model():
    X, y = make_classification(
        n_samples=1200, n_features=N_FEATURES, n_informative=5,
        n_redundant=2, weights=[0.85, 0.15], flip_y=0.02,
        class_sep=1.1, random_state=SEED)
    Xtr, Xte, ytr, yte = train_test_split(
        X, y, test_size=0.25, stratify=y, random_state=SEED)
    pipe = Pipeline([("scale", StandardScaler()),
                     ("clf", GradientBoostingClassifier(random_state=SEED))])
    pipe.fit(Xtr, ytr)
    return pipe, Xte, yte


_PORT = None


def call(path, payload=None, method="GET"):
    url = f"http://{HOST}:{_PORT}{path}"
    data = json.dumps(payload).encode() if payload is not None else None
    req = urllib.request.Request(
        url, data=data, method=method,
        headers={"Content-Type": "application/json"} if data else {})
    try:
        with urllib.request.urlopen(req, timeout=10) as resp:
            return resp.status, json.loads(resp.read())
    except urllib.error.HTTPError as exc:
        return exc.code, json.loads(exc.read())


def percentile(values, p):
    s = sorted(values)
    k = (len(s) - 1) * p / 100
    lo, hi = int(k), min(int(k) + 1, len(s) - 1)
    return s[lo] + (s[hi] - s[lo]) * (k - lo)


def main():
    global _MODEL, _PORT
    print("  Experiment 15 -- a model deployed as a REST endpoint, "
          "actually served")

    # Step 1: Train the model, and save it
    model, X_test, y_test = build_model()
    path = os.path.join(tempfile.gettempdir(), "cloud13b_endpoint.joblib")
    joblib.dump(model, path)
    _MODEL = joblib.load(path)
    print(f"\n    artefact loaded from disk: {os.path.getsize(path):,} bytes")

    # Step 2: Serve it
    HTTPServer.allow_reuse_address = True
    server = HTTPServer((HOST, PORT), Handler)
    global _PORT
    _PORT = server.server_address[1]
    thread = threading.Thread(target=server.serve_forever, daemon=True)
    thread.start()
    # [Changed: this printed the port, which the system picks afresh on every run.]
    print(f"    endpoint listening on http://{HOST}, on a port the system chose  "
          f"(a REAL HTTP server)")

    try:
        # Step 3: Check its health
        code, body = call("/ping")
        print(f"\n    GET /ping        -> {code} {body}")
        assert code == 200 and body["model_loaded"] is True
        print("""         /ping answers WITHOUT running the model. A health check
         that does real inference marks the container unhealthy
         whenever the model is merely slow, and the platform then
         kills a container that was working -- a self-inflicted
         outage, and a classic one""")

        # Step 4: Ask for a prediction
        sample = X_test[:3].tolist()
        code, body = call("/invocations", {"instances": sample}, "POST")
        print(f"\n    POST /invocations with 3 rows -> {code}")
        print(f"      predictions   : {body['predictions']}")
        print(f"      probabilities : {body['probabilities']}")
        assert code == 200 and len(body["predictions"]) == 3
        local = model.predict(X_test[:3]).tolist()
        assert body["predictions"] == local
        print("""         the endpoint's answers are IDENTICAL to calling the model
         in-process. That equality is the deployment test worth
         writing: serving must not change predictions, and a
         preprocessing step that lives in your notebook rather than
         in the pipeline is exactly how it does""")

        # Step 5: Send a batch
        code, body = call("/invocations",
                          {"instances": X_test.tolist()}, "POST")
        assert code == 200
        preds = np.array(body["predictions"])
        acc = (preds == y_test).mean()
        print(f"\n    POST /invocations with all {len(X_test)} rows -> {code}, "
              f"accuracy {acc:.4f}")
        assert acc > 0.90

        # Step 6: Send bad requests
        print("\n    error handling, which is most of a real endpoint:")
        cases = [
            ("wrong feature count", {"instances": [[1.0, 2.0]]}, "POST",
             "/invocations"),
            ("not a list", {"instances": "hello"}, "POST", "/invocations"),
            ("empty body", {}, "POST", "/invocations"),
            ("wrong route", {"instances": sample}, "POST", "/predict"),
            ("wrong route, GET", None, "GET", "/predict"),
        ]
        print(f"      {'case':<24}{'status':>8}  message")
        for label, payload, method, route in cases:
            code, body = call(route, payload, method)
            msg = body.get("error", "")[:46]
            print(f"      {label:<24}{code:>8}  {msg}")
            assert 400 <= code < 500, "a client mistake must not be a 5xx"
        print("""         EVERY ONE IS A 4XX, NOT A 5XX, and that distinction is
         operational rather than pedantic: 5xx means YOUR service is
         broken and should page someone. If malformed client input
         returns 500, your error alarm fires for other people's bugs
         and you stop trusting it""")

        # Step 7: Measure the latency
        print("\n    latency over 200 single-row requests:")
        _STATE["latencies"].clear()
        for i in range(200):
            code, _ = call("/invocations",
                           {"instances": [X_test[i % len(X_test)].tolist()]},
                           "POST")
            assert code == 200
        lat = _STATE["latencies"]
        p50, p95, p99 = (percentile(lat, p) for p in (50, 95, 99))
        mean = sum(lat) / len(lat)
        print(f"      mean {mean:.3f} ms   p50 {p50:.3f} ms   "
              f"p95 {p95:.3f} ms   p99 {p99:.3f} ms")
        assert p99 >= p50
        print(f"""         p99 is {p99 / p50:.1f}x p50 on an idle laptop serving one model.
         On a shared endpoint under load that ratio grows, which is
         why the alarm in experiment 13 is on p99 and not the mean.
         These are SERVER-SIDE numbers; a client also pays network
         time, and the user's experience is the sum""")

        # Step 8: Batch, against one at a time
        print("\n    one request of 100 rows against 100 requests of one row:")
        _STATE["latencies"].clear()
        t0 = time.perf_counter()
        call("/invocations", {"instances": X_test[:100].tolist()}, "POST")
        batched = (time.perf_counter() - t0) * 1000
        t0 = time.perf_counter()
        for i in range(100):
            call("/invocations", {"instances": [X_test[i].tolist()]}, "POST")
        singly = (time.perf_counter() - t0) * 1000
        print(f"      batched  : {batched:8.2f} ms total")
        print(f"      one by one: {singly:8.2f} ms total "
              f"({singly / batched:.0f}x)")
        assert singly > batched
        print(f"""         {singly / batched:.0f}x, and none of it is the model -- it is per-request
         overhead: HTTP, JSON parsing, and a NumPy call whose fixed
         cost is paid 100 times instead of once.
         This is why batch transform exists alongside real-time
         endpoints. If you are scoring a file of a million rows,
         calling an endpoint a million times is the expensive way to
         do arithmetic""")

        # Step 9: Read the metrics
        code, metrics = call("/metrics")
        print(f"\n    GET /metrics -> {metrics['invocations']:,} invocations, "
              f"{metrics['errors_4xx']} 4xx, {metrics['errors_5xx']} 5xx")
        assert metrics["errors_5xx"] == 0

    finally:
        server.shutdown()
        server.server_close()
        os.remove(path)
        print("\n    endpoint shut down and artefact removed.")

    # Step 10: Compare with a managed endpoint
    print("\n    what SageMaker adds that this server does not have:")
    print(f"      {'':<26}{'this script':<22}{'a managed endpoint'}")
    for label, here, cloud in (
            ("TLS", "no", "yes, terminated for you"),
            ("authentication", "NONE -- anyone", "IAM-signed requests"),
            ("load balancing", "one process", "across instances and AZs"),
            ("autoscaling", "no", "on InvocationsPerInstance"),
            ("blue/green deploy", "no", "traffic shifted gradually"),
            ("metrics", "the dict above", "CloudWatch, automatically"),
            ("cost", "electricity", "PER HOUR, until deleted")):
        print(f"      {label:<26}{here:<22}{cloud}")
    hourly = f.EC2["m5.large"]
    print(f"\n      an ml.m5.large endpoint: ${hourly:.4f}/hour "
          f"= ${hourly * f.HOURS_PER_MONTH:,.2f}/month, called or not")
    print("""         THE LAST ROW AGAIN. A training job stops; an endpoint does
         not. Deleting the endpoint is a step in the experiment, not
         an afterthought -- and if traffic is occasional, a serverless
         endpoint or batch transform costs a fraction of it""")


if __name__ == "__main__":
    main()

5. Execution and Results

The procedure, on the console and CLI, 15_deploy.md:

NOT RUN HERE

15_deploy.md was not run: it needs a cloud account, and this repository has none. Nothing on this page claims an output it did not produce.

The Python model, which runs, 15_deploy_endpoint.py, for experiment 15:

OUTPUT

  Experiment 15 -- a model deployed as a REST endpoint, actually served

    artefact loaded from disk: 138,945 bytes
    endpoint listening on http://127.0.0.1, on a port the system chose  (a REAL HTTP server)

    GET /ping        -> 200 {'status': 'healthy', 'model_loaded': True}
         /ping answers WITHOUT running the model. A health check
         that does real inference marks the container unhealthy
         whenever the model is merely slow, and the platform then
         kills a container that was working -- a self-inflicted
         outage, and a classic one

    POST /invocations with 3 rows -> 200
      predictions   : [0, 0, 0]
      probabilities : [0.015383, 0.008643, 0.014492]
         the endpoint's answers are IDENTICAL to calling the model
         in-process. That equality is the deployment test worth
         writing: serving must not change predictions, and a
         preprocessing step that lives in your notebook rather than
         in the pipeline is exactly how it does

    POST /invocations with all 300 rows -> 200, accuracy 0.9467

    error handling, which is most of a real endpoint:
      case                      status  message
      wrong feature count          400  each instance needs 10 features, got shape [1,
      not a list                   400  body must be {'instances': [[...], ...]}
      empty body                   400  body must be {'instances': [[...], ...]}
      wrong route                  404  not found
      wrong route, GET             404  not found
         EVERY ONE IS A 4XX, NOT A 5XX, and that distinction is
         operational rather than pedantic: 5xx means YOUR service is
         broken and should page someone. If malformed client input
         returns 500, your error alarm fires for other people's bugs
         and you stop trusting it

    latency over 200 single-row requests:
      mean 0.735 ms   p50 0.709 ms   p95 0.972 ms   p99 1.184 ms
         p99 is 1.7x p50 on an idle laptop serving one model.
         On a shared endpoint under load that ratio grows, which is
         why the alarm in experiment 13 is on p99 and not the mean.
         These are SERVER-SIDE numbers; a client also pays network
         time, and the user's experience is the sum

    one request of 100 rows against 100 requests of one row:
      batched  :     2.74 ms total
      one by one:   155.69 ms total (57x)
         57x, and none of it is the model -- it is per-request
         overhead: HTTP, JSON parsing, and a NumPy call whose fixed
         cost is paid 100 times instead of once.
         This is why batch transform exists alongside real-time
         endpoints. If you are scoring a file of a million rows,
         calling an endpoint a million times is the expensive way to
         do arithmetic

    GET /metrics -> 703 invocations, 5 4xx, 0 5xx

    endpoint shut down and artefact removed.

    what SageMaker adds that this server does not have:
                                this script           a managed endpoint
      TLS                       no                    yes, terminated for you
      authentication            NONE -- anyone        IAM-signed requests
      load balancing            one process           across instances and AZs
      autoscaling               no                    on InvocationsPerInstance
      blue/green deploy         no                    traffic shifted gradually
      metrics                   the dict above        CloudWatch, automatically
      cost                      electricity           PER HOUR, until deleted

      an ml.m5.large endpoint: $0.0960/hour = $70.08/month, called or not
         THE LAST ROW AGAIN. A training job stops; an endpoint does
         not. Deleting the endpoint is a step in the experiment, not
         an afterthought -- and if traffic is occasional, a serverless
         endpoint or batch transform costs a fraction of it

/PING MUST NOT RUN THE MODEL

A health check that does real inference marks the container unhealthy whenever the model is merely slow — and the platform then kills a container that was working. A self-inflicted outage, and a classic one.

Error handling, which is most of a real endpoint:

Case Status
wrong feature count 400
body not a list 400
empty body 400
wrong route (POST and GET) 404

Every one is a 4xx, not a 5xx, and the distinction is operational rather than pedantic: 5xx means your service is broken and should page someone. If malformed client input returns 500, your error alarm fires for other people's bugs and you stop trusting it.

Latency, over 200 real requests, and the batching result, are the timed lines above: they measure this machine at one moment and differ from run to run. What does not change is their shape — p99 above p50 even on an idle machine serving one model, and one request of 100 rows far faster than 100 requests of one row, none of the difference being the model: it is per-request overhead, HTTP, JSON parsing, and a NumPy call whose fixed cost is paid 100 times instead of once. Corrected: this page gave "p99 is 1.8× p50" and "37×" as if fixed; they were one run's figures, and three runs made here gave batching factors of 47×, 57× and 68×. Under load the p99 ratio grows — which is why experiment 13's alarm is on p99.

NOTE

If you are scoring a million rows, calling an endpoint a million times is the expensive way to do arithmetic.

What SageMaker adds that this server does not have:

This script A managed endpoint
TLS no terminated for you
Authentication NONE — anyone IAM-signed requests
Load balancing one process across instances and AZs
Autoscaling no on InvocationsPerInstance
Blue/green no traffic shifted gradually
Cost electricity $70/month, called or not

Deleting the endpoint is a step in the experiment, not an afterthought.

Changed: the program printed the port it listened on, which the system picks afresh on every run; it now says so.

RESULT

The endpoint's predictions equal the model's in-process; bad input gets a 4xx, never a 5xx; batching beat one row per request many times over.


What the runner asserts

Script Experiments Real?
01_vm_and_hosting.py 1, 2, 7 a real web server; overcommit modelled
03_iam_and_account.py 3, 10 the real IAM algorithm
04_storage.py 4, 5, 6 real key semantics, real arithmetic
09_etl_warehouse.py 8, 9, 12 real SQLite → real DuckDB
11_train_and_automl.py 11, 14 a real model, a real 25-fit search
13_monitoring_autoscale.py 13 a real control loop
15_deploy_endpoint.py 15 a real HTTP endpoint, called over TCP

Plus the audit: 14 Markdown files, every one carrying NOT EXECUTED, each naming the service it needs.


Lab examination

Two hours on a console, one experiment number, then a viva.

What costs marks:

What earns them:

Written-out instructions

These experiments are console procedures rather than programs, so each one is written out as a page.

EXPERIMENT 1

Create a virtual machine in VMware Workstation

EXPERIMENT 2

Install and configure Apache/XAMPP on the VM and host a page

EXPERIMENT 3

Create and configure a cloud account (AWS/Azure/GCP free tier)

EXPERIMENT 4

Create and manage storage buckets; upload and access datasets

EXPERIMENT 5

Launch an instance and configure block storage (EBS)

EXPERIMENT 6

Create and configure file storage on a cloud VM (EFS)

EXPERIMENT 7

Set up Jupyter Notebook / Colab on a cloud VM

EXPERIMENT 8

Connect to cloud-hosted database services (RDS, BigQuery, Cosmos DB)

EXPERIMENT 10

Launch a SageMaker notebook, attach an IAM role and an S3 bucket

EXPERIMENT 11

Build a classification/regression model on a managed ML platform

EXPERIMENT 12

A simple ETL job: extract, transform, load into a cloud warehouse

EXPERIMENT 13

Use CloudWatch/Stackdriver to monitor endpoints, set alarms and auto-scale

EXPERIMENT 14

Use cloud AutoML services for a dataset prediction task

EXPERIMENT 15

Deploy a trained ML model as a REST API endpoint

Each program, on its own page

The same experiments, one page each, so a program can be reached by what it does rather than by its number.

RUNS

A virtual machine, a web server on it in the Cloud

RUNS

Cloud account setup, and IAM roles for SageMaker

RUNS

Object, block and file storage on the cloud

RUNS

Cloud databases, a batch ETL pipeline

RUNS

Build a model on a managed ML platform in the Cloud

RUNS

CloudWatch/Stackdriver: monitor an endpoint, set alarms

RUNS

Deploy a trained ML model as a REST API endpoint in the Cloud