Skip to the content
On this page
  1. Read this before you read anything else
  2. Experiment 1 — Installation and setup of a Hadoop single-node cluster
  3. Experiment 2 — Hadoop's directory structure and basic commands
  4. Experiment 3 — Hadoop's architecture, from its logs
  5. Experiment 4 — Storing a large file in HDFS: blocks and replication
  6. Experiment 5 — Fault tolerance and recovery
  7. Experiment 6 — YARN: the ResourceManager, the NodeManager and the schedulers
  8. Experiment 7 — Word count in MapReduce
  9. Experiment 8 — An inverted index in MapReduce
  10. Experiment 9 — Data analysis with Pig Latin
  11. Experiment 10 — Hive: tables, partitions and buckets
  12. Experiment 11 — Importing from a relational database with Sqoop
  13. Experiment 12 — Capturing log data with Flume
  14. Experiment 13 — Avro and Parquet
  15. Experiment 14 — An end-to-end ingestion workflow, batch and streaming
  16. Experiment 15 — HBase: tables and CRUD
  17. Experiment 16 — Coordination with ZooKeeper
  18. Experiment 17 — Spark over HBase
  19. What the runner asserts
  20. Lab examination
  21. Each program, on its own page

17 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.

Code lives in labs/course-12b-bigdata/.

Read this before you read anything else

Every file in this lab runs, on a real Hadoop 3.3.6 cluster. Most experiments have two halves:

Half Files Status
The tool you run on a cluster 15 files — shell scripts of hdfs, yarn, sqoop and zkCli.sh commands, Pig, HiveQL, the HBase shell, a Flume agent, Java and Scala Run, each on a fresh cluster, by tools/data-science/hadoop_lab.py; what it printed is under 5. Execution and Results
The check 14 programs Executed and asserted by tools/data-science/run_bigdata_labs.py, which also runs the tool files and checks their answers

(Updated October 2026: Hadoop, Pig, Hive, HBase, ZooKeeper, Sqoop and Flume could not be installed where these labs are checked, and every tool file said NOT EXECUTED. They now install from archive.apache.org, with MariaDB from the Ubuntu archive for Sqoop, and running the files found faults that reading them had not — a Pig keyword used as a field name, two Hive statements that stop the script, a Sqoop query with two columns of one name, a Flume agent that dropped its HDFS sink and routed nothing, an HBase filter that let rows through, a Spark scan that saw no rows. Each is corrected in its file and noted under its experiment.)

bash tools/data-science/setup_hadoop.sh         # Hadoop, Pig, Hive, HBase, ZooKeeper, Sqoop, Flume
bash tools/data-science/setup_spark.sh          # PySpark, for experiment 17
pip install -r tools/requirements.txt
python3 tools/data-science/run_bigdata_labs.py  # about half an hour; --audit-only skips the cluster

HOW THE TOOL FILES WERE RUN

hadoop_lab.py formats and starts a cluster in a temporary folder for each file — a NameNode, four DataNodes (so replication 3 can survive the loss of one, and a block can be re-replicated), a SecondaryNameNode, a ResourceManager, a NodeManager and the JobHistory server — then runs the file as a student types it, printing $ command before each command, and stops everything at the end.

Where a file assumes something is already there — the sales file in HDFS, a database for Sqoop, HBase running — a _drive_ script beside it does that first, and those commands are printed too, so the transcript shows everything that ran. One setting differs from a default cluster, and the experiment that depends on it says so: a DataNode is declared dead after 60 s, not 630 s.

A cluster makes new ids, ports and times on every run — block ids, application ids, the port a DataNode took, how long a job ran. The output shown is one run's, unchanged; capture_lab_outputs.py --check compares a rerun with those values masked, and every other line must repeat exactly.

The cross-course check

Experiments 10, 14 and 17 all use Business Intelligence Tools' star schema, imported rather than copied. fixtures.py loads labs/course-11-bi/fixtures.py by path at import time, so the two courses cannot drift.

South = ₹10,360 is produced by Business Intelligence Tools' DAX CALCULATE, by DuckDB and by Hive in experiment 10, and by Spark — reading HBase — in experiment 17. If they ever disagree, one of them is wrong and verify_all.sh says so.


Experiment 1 — Installation and setup of a Hadoop single-node cluster

1. Question

Install Hadoop on one machine, configure it, and start its daemons.

2. Aim

Download Hadoop, write its four configuration files, format HDFS, start the five daemons and check each one is up.

3. Steps

On the cluster, 01_install_hadoop.sh:

  1. Check the prerequisites.
  2. Download and unpack Hadoop.
  3. Write the four configuration files.
  4. Format HDFS, and start the five daemons.
  5. Stop them.

THE THREE INSTALLATION FAILURES EVERYONE HITS

jps should show five processes: NameNode, DataNode, SecondaryNameNode, ResourceManager, NodeManager.

4. Programme

On the cluster, 01_install_hadoop.sh:

# Experiment 1 -- installation and setup of a Hadoop single-node cluster
#
# Run it: bash 01_install_hadoop.sh, on Linux with Java 11 installed. It was run where these
# labs are checked, in an empty home directory, and the lab page shows what it printed.
# [Changed: this said the file had never been run, as Hadoop could not be installed there. It
# installs from archive.apache.org, and the corrections below are what running it found.]
#
# The runnable half is none -- installation has no query logic to verify
#
# Step 1: Check the prerequisites
java -version                       # Hadoop 3.x needs Java 8 or 11

# On your own machine, run Hadoop as its own user, with passwordless ssh to
# localhost -- start-dfs.sh and start-yarn.sh ssh to every host they start a
# daemon on, even in "single node":
#   sudo adduser hadoop && su - hadoop
#   ssh-keygen -t rsa -P '' -f ~/.ssh/id_rsa
#   cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
#   chmod 600 ~/.ssh/authorized_keys
#   ssh localhost                   # must succeed WITHOUT a password
# Where this was run there is no ssh server, so the daemons are started one by
# one below, with the same commands start-dfs.sh runs on each host.

# Step 2: Download and unpack Hadoop
[ -f hadoop-3.3.6.tar.gz ] || \
  curl -fO https://archive.apache.org/dist/hadoop/common/hadoop-3.3.6/hadoop-3.3.6.tar.gz
tar -xzf hadoop-3.3.6.tar.gz && mv hadoop-3.3.6 ~/hadoop
# [Corrected: the download was from dlcdn.apache.org, which keeps only the
# current releases -- 3.3.6 is no longer there, and wget got a 404. Every
# release stays in archive.apache.org. And it was moved to /usr/local/hadoop
# with sudo; your home directory needs no sudo.]

cat >> ~/.bashrc <<'EOF'
export HADOOP_HOME=~/hadoop
export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64
export PATH=$PATH:$HADOOP_HOME/bin:$HADOOP_HOME/sbin
export HADOOP_CONF_DIR=$HADOOP_HOME/etc/hadoop
EOF
source ~/.bashrc

# Step 3: Write the four configuration files
cat > $HADOOP_CONF_DIR/core-site.xml <<'EOF'
<configuration>
  <property><name>fs.defaultFS</name><value>hdfs://localhost:9000</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/hdfs-site.xml <<EOF
<configuration>
  <!-- 1, not 3: there is only one node -->
  <property><name>dfs.replication</name><value>1</value></property>
  <property><name>dfs.namenode.name.dir</name><value>file://$HOME/hadoop_store/hdfs/namenode</value></property>
  <property><name>dfs.datanode.data.dir</name><value>file://$HOME/hadoop_store/hdfs/datanode</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/mapred-site.xml <<EOF
<configuration>
  <property><name>mapreduce.framework.name</name><value>yarn</value></property>
  <property><name>yarn.app.mapreduce.am.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
  <property><name>mapreduce.map.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
  <property><name>mapreduce.reduce.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/yarn-site.xml <<'EOF'
<configuration>
  <property><name>yarn.nodemanager.aux-services</name><value>mapreduce_shuffle</value></property>
</configuration>
EOF
echo "export JAVA_HOME=$JAVA_HOME" >> $HADOOP_CONF_DIR/hadoop-env.sh
# [Corrected: the four files were listed as property names only; they are
# written out here, under your home directory rather than /usr/local, which
# needs no sudo. Hadoop 3 also needs HADOOP_MAPRED_HOME passed to MapReduce's
# containers, or every job fails with "Could not find or load main class
# org.apache.hadoop.mapreduce.v2.app.MRAppMaster".]

# Step 4: Format HDFS, and start the five daemons
hdfs namenode -format 2>&1 | grep "has been successfully formatted"
#   ONCE. Re-formatting destroys the cluster. (It prints a hundred log lines;
#   this is the one that says it worked.)
hdfs --daemon start namenode
hdfs --daemon start datanode
hdfs --daemon start secondarynamenode
yarn --daemon start resourcemanager
yarn --daemon start nodemanager
#   = start-dfs.sh and start-yarn.sh, which run exactly these on each host,
#     over ssh

sleep 15                            # the DataNode and NodeManager register
jps | awk '{print $2}' | sort       # expect: NameNode, DataNode,
                                    # SecondaryNameNode, ResourceManager,
                                    # NodeManager  -- five processes, and Jps
hdfs dfsadmin -report 2>/dev/null | grep "Live datanodes"
hdfs dfs -mkdir -p /user/$USER && hdfs dfs -ls /user
# web UIs -- 200 means each is up:
curl -sL -o /dev/null -w "NameNode        http://localhost:9870  %{http_code}\n" http://localhost:9870
curl -s -o /dev/null -w "ResourceManager http://localhost:8088  %{http_code}\n" http://localhost:8088/cluster

# Step 5: Stop them
yarn --daemon stop nodemanager
yarn --daemon stop resourcemanager
hdfs --daemon stop secondarynamenode
hdfs --daemon stop datanode
hdfs --daemon stop namenode
#   = stop-yarn.sh and stop-dfs.sh

# --- the three failures everyone hits --------------------------------------
# 1. JAVA_HOME not set INSIDE hadoop-env.sh (the shell export is not enough)
# 2. re-running `hdfs namenode -format` after storing data: the DataNode's
#    clusterID no longer matches the NameNode's, and the DataNode will not
#    start. Fix: delete the datanode directory, or edit its VERSION file.
# 3. ssh localhost prompting for a password -- start-dfs.sh hangs for ever

5. Execution and Results

On the cluster, 01_install_hadoop.sh:

OUTPUT

$ java -version
openjdk version "11.0.32.1" 2026-08-18
OpenJDK Runtime Environment (build 11.0.32.1+1-post-1ubuntu1-24.04-Ubuntu)
OpenJDK 64-Bit Server VM (build 11.0.32.1+1-post-1ubuntu1-24.04-Ubuntu, mixed mode, sharing)
$ [ -f hadoop-3.3.6.tar.gz ] || \
    curl -fO https://archive.apache.org/dist/hadoop/common/hadoop-3.3.6/hadoop-3.3.6.tar.gz
$ tar -xzf hadoop-3.3.6.tar.gz && mv hadoop-3.3.6 ~/hadoop
$ cat >> ~/.bashrc <<'EOF'
  export HADOOP_HOME=~/hadoop
  export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64
  export PATH=$PATH:$HADOOP_HOME/bin:$HADOOP_HOME/sbin
  export HADOOP_CONF_DIR=$HADOOP_HOME/etc/hadoop
  EOF
$ source ~/.bashrc
$ cat > $HADOOP_CONF_DIR/core-site.xml <<'EOF'
  <configuration>
    <property><name>fs.defaultFS</name><value>hdfs://localhost:9000</value></property>
  </configuration>
  EOF
$ cat > $HADOOP_CONF_DIR/hdfs-site.xml <<EOF
  <configuration>
    <!-- 1, not 3: there is only one node -->
    <property><name>dfs.replication</name><value>1</value></property>
    <property><name>dfs.namenode.name.dir</name><value>file://$HOME/hadoop_store/hdfs/namenode</value></property>
    <property><name>dfs.datanode.data.dir</name><value>file://$HOME/hadoop_store/hdfs/datanode</value></property>
  </configuration>
  EOF
$ cat > $HADOOP_CONF_DIR/mapred-site.xml <<EOF
  <configuration>
    <property><name>mapreduce.framework.name</name><value>yarn</value></property>
    <property><name>yarn.app.mapreduce.am.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
    <property><name>mapreduce.map.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
    <property><name>mapreduce.reduce.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
  </configuration>
  EOF
$ cat > $HADOOP_CONF_DIR/yarn-site.xml <<'EOF'
  <configuration>
    <property><name>yarn.nodemanager.aux-services</name><value>mapreduce_shuffle</value></property>
  </configuration>
  EOF
$ echo "export JAVA_HOME=$JAVA_HOME" >> $HADOOP_CONF_DIR/hadoop-env.sh
$ hdfs namenode -format 2>&1 | grep "has been successfully formatted"
2026-10-04 23:03:54,977 INFO common.Storage: Storage directory ~/hadoop_store/hdfs/namenode has been successfully formatted.
$ hdfs --daemon start namenode
$ hdfs --daemon start datanode
$ hdfs --daemon start secondarynamenode
$ yarn --daemon start resourcemanager
$ yarn --daemon start nodemanager
$ sleep 15
$ jps | awk '{print $2}' | sort
DataNode
Jps
NameNode
NodeManager
ResourceManager
SecondaryNameNode
$ hdfs dfsadmin -report 2>/dev/null | grep "Live datanodes"
Live datanodes (1):
$ hdfs dfs -mkdir -p /user/$USER && hdfs dfs -ls /user
Found 1 items
drwxr-xr-x   - root supergroup          0 2026-10-04 23:04 /user/root
$ curl -sL -o /dev/null -w "NameNode        http://localhost:9870  %{http_code}\n" http://localhost:9870
NameNode        http://localhost:9870  200
$ curl -s -o /dev/null -w "ResourceManager http://localhost:8088  %{http_code}\n" http://localhost:8088/cluster
ResourceManager http://localhost:8088  200
$ yarn --daemon stop nodemanager
$ yarn --daemon stop resourcemanager
$ hdfs --daemon stop secondarynamenode
$ hdfs --daemon stop datanode
$ hdfs --daemon stop namenode

This one runs alone, not on the lab's cluster: _drive_01_install_hadoop.py gives it an empty home directory and the Hadoop tarball, and the script does everything else — unpacks, configures, formats, starts, checks and stops. Where it was run there is no ssh server, so the daemons are started one by one with hdfs --daemon start, the command start-dfs.sh runs on each host over ssh.

RESULT

Hadoop 3.3.6 installed in an empty home directory, HDFS formatted, the five daemons running — jps lists all five — one live DataNode, and both web interfaces answering 200. Two corrections: the download moved to archive.apache.org, and MapReduce needs HADOOP_MAPRED_HOME passed to its containers.

Experiment 2 — Hadoop's directory structure and basic commands

1. Question

Explore the Hadoop directory structure and its basic commands — the hadoop fs operations.

2. Aim

Make, list, copy, move and remove files in HDFS; count the space they use; set permissions and replication; and ask where a file's blocks are.

3. Steps

On the cluster, 02_hdfs_commands.sh:

  1. Make a directory, and list it.
  2. Put a file in, and read it back.
  3. Copy, move and remove.
  4. Count the space used.
  5. Set permissions and replication.
  6. Ask where the blocks are, and how the cluster is.

THE FOUR HADOOP FS FACTS WORTH MARKS

4. Programme

On the cluster, 02_hdfs_commands.sh:

# Experiment 2 -- explore the Hadoop directory structure and basic hadoop fs commands
#
# Run it: bash 02_hdfs_commands.sh, with a cluster running (experiment 1). It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is none -- these are filesystem commands, verified by the arithmetic in 04
#
# Step 1: Make a directory, and list it
hdfs dfs -mkdir -p /user/student/sales
hdfs dfs -ls /user/student
hdfs dfs -ls -R /user                 # recursive
# Step 2: Put a file in, and read it back
hdfs dfs -put sales.csv /user/student/sales/
hdfs dfs -cat /user/student/sales/sales.csv | head
hdfs dfs -tail /user/student/sales/sales.csv
hdfs dfs -get /user/student/sales/sales.csv ./back.csv
# Step 3: Copy, move and remove
hdfs dfs -cp  /user/student/sales/sales.csv /user/student/copy.csv
hdfs dfs -mv  /user/student/copy.csv /user/student/moved.csv
hdfs dfs -rm -r /user/student/sales   # goes to .Trash, not to nothing
# Step 4: Count the space used
hdfs dfs -du -h /user/student
hdfs dfs -df -h /
hdfs dfs -count /user/student         # DIRS FILES BYTES
# Step 5: Set permissions and replication
hdfs dfs -chmod 640 /user/student/moved.csv
hdfs dfs -chown student:analysts /user/student/moved.csv
hdfs dfs -setrep -w 2 /user/student/moved.csv     # change replication
hdfs dfs -stat "%r %o %b" /user/student/moved.csv # replication blocksize bytes

# Step 6: Ask where the blocks are, and how the cluster is
hdfs fsck /user/student -files -blocks -locations  # WHERE each block lives
hdfs dfsadmin -report                              # per-DataNode capacity
hdfs dfsadmin -safemode get                        # ON during startup

# --- the four things that surprise people ----------------------------------
# 1. `hadoop fs` and `hdfs dfs` are the same command. `hadoop fs` also works
#    on local and S3 paths; `hdfs dfs` is HDFS only.
# 2. THERE IS NO `cd`. HDFS has no working directory -- every path is
#    absolute, or relative to /user/$USER.
# 3. `-rm` moves to .Trash and still costs quota for a day. Use -skipTrash
#    when you mean it.
# 4. There is no in-place edit. HDFS is WRITE-ONCE, APPEND-ONLY: to change one
#    byte you rewrite the file. That single constraint is why HDFS can drop
#    file locking, and why it suits analytics and not OLTP.

5. Execution and Results

On the cluster, 02_hdfs_commands.sh:

OUTPUT

$ hdfs dfs -mkdir -p /user/student/sales
$ hdfs dfs -ls /user/student
Found 1 items
drwxr-xr-x   - root supergroup          0 2026-10-04 23:05 /user/student/sales
$ hdfs dfs -ls -R /user
drwxr-xr-x   - root supergroup          0 2026-10-04 23:05 /user/student
drwxr-xr-x   - root supergroup          0 2026-10-04 23:05 /user/student/sales
$ hdfs dfs -put sales.csv /user/student/sales/
$ hdfs dfs -cat /user/student/sales/sales.csv | head
date_key,store,region,product,category,qty,list_price
D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0
D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0
D1,Guntur,South,Tea 500g,Grocery,8,210.0
D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0
D2,Hyderabad,North,Notebook,Stationery,20,40.0
D3,Guntur,South,Tea 500g,Grocery,12,210.0
D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0
D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0
D4,Hyderabad,North,Notebook,Stationery,15,40.0
$ hdfs dfs -tail /user/student/sales/sales.csv
date_key,store,region,product,category,qty,list_price
D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0
D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0
D1,Guntur,South,Tea 500g,Grocery,8,210.0
D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0
D2,Hyderabad,North,Notebook,Stationery,20,40.0
D3,Guntur,South,Tea 500g,Grocery,12,210.0
D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0
D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0
D4,Hyderabad,North,Notebook,Stationery,15,40.0
$ hdfs dfs -get /user/student/sales/sales.csv ./back.csv
$ hdfs dfs -cp  /user/student/sales/sales.csv /user/student/copy.csv
$ hdfs dfs -mv  /user/student/copy.csv /user/student/moved.csv
$ hdfs dfs -rm -r /user/student/sales
2026-10-04 23:06:06,486 INFO fs.TrashPolicyDefault: Moved: 'hdfs://localhost:9000/user/student/sales' to trash at: hdfs://localhost:9000/user/root/.Trash/Current/user/student/sales
$ hdfs dfs -du -h /user/student
468  1.4 K  /user/student/moved.csv
$ hdfs dfs -df -h /
Filesystem                 Size    Used  Available  Use%
hdfs://localhost:9000  1007.9 G  98.8 K     36.1 G    0%
$ hdfs dfs -count /user/student
           1            1                468 /user/student
$ hdfs dfs -chmod 640 /user/student/moved.csv
$ hdfs dfs -chown student:analysts /user/student/moved.csv
$ hdfs dfs -setrep -w 2 /user/student/moved.csv
Replication 2 set: /user/student/moved.csv
Waiting for /user/student/moved.csv ...
WARNING: the waiting time may be long for DECREASING the number of replications.
. done
$ hdfs dfs -stat "%r %o %b" /user/student/moved.csv
2 134217728 468
$ hdfs fsck /user/student -files -blocks -locations
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&locations=1&path=%2Fuser%2Fstudent
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student at Sun Oct 04 23:06:30 UTC 2026

/user/student <dir>
/user/student/moved.csv 468 bytes, replicated: replication=2, 1 block(s):  OK
0. BP-1985187428-127.0.0.1-1791155116848:blk_1073741826_1002 len=468 Live_repl=2  [DatanodeInfoWithStorage[127.0.0.1:9886,DS-aeb1c4f1-1445-4508-8b83-5f8852f1b202,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-c7bba049-11ba-4c05-a5e4-c385376f29cb,DISK]]


Status: HEALTHY
 Number of data-nodes:  4
 Number of racks:       1
 Total dirs:            1
 Total symlinks:        0

Replicated Blocks:
 Total size:    468 B
 Total files:   1
 Total blocks (validated):  1 (avg. block size 468 B)
 Minimally replicated blocks:   1 (100.0 %)
 Over-replicated blocks:    0 (0.0 %)
 Under-replicated blocks:   0 (0.0 %)
 Mis-replicated blocks:     0 (0.0 %)
 Default replication factor:    3
 Average block replication: 2.0
 Missing blocks:        0
 Corrupt blocks:        0
 Missing replicas:      0 (0.0 %)
 Blocks queued for replication: 0

Erasure Coded Block Groups:
 Total size:    0 B
 Total files:   0
 Total block groups (validated):    0
 Minimally erasure-coded block groups:  0
 Over-erasure-coded block groups:   0
 Under-erasure-coded block groups:  0
 Unsatisfactory placement block groups: 0
 Average block group size:  0.0
 Missing block groups:      0
 Corrupt block groups:      0
 Missing internal blocks:   0
 Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:06:30 UTC 2026 in 9 milliseconds


The filesystem under path '/user/student' is HEALTHY
$ hdfs dfsadmin -report
Configured Capacity: 1082212696064 (1007.89 GB)
Present Capacity: 38770710875 (36.11 GB)
DFS Remaining: 38770610176 (36.11 GB)
DFS Used: 100699 (98.34 KB)
DFS Used%: 0.00%
Replicated Blocks:
    Under replicated blocks: 0
    Blocks with corrupt replicas: 0
    Missing blocks: 0
    Missing blocks (with replication factor 1): 0
    Low redundancy blocks with highest priority to recover: 0
    Pending deletion blocks: 0
Erasure Coded Block Groups:
    Low redundancy block groups: 0
    Block groups with corrupt internal blocks: 0
    Missing block groups: 0
    Low redundancy blocks with highest priority to recover: 0
    Pending deletion blocks: 0

-------------------------------------------------
Live datanodes (4):

Name: 127.0.0.1:9866 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088212001 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:22 UTC 2026
Num of Blocks: 1


Name: 127.0.0.1:9886 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088212001 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:25 UTC 2026
Num of Blocks: 1


Name: 127.0.0.1:9896 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25534 (24.94 KB)
Non DFS Used: 30088211522 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:28 UTC 2026
Num of Blocks: 2


Name: 127.0.0.1:9906 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088207905 (28.02 GB)
DFS Remaining: 9692655616 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:30 UTC 2026
Last Block Report: Sun Oct 04 23:05:30 UTC 2026
Num of Blocks: 1


$ hdfs dfsadmin -safemode get
Safe mode is OFF

RESULT

Every hdfs dfs operation ran against the cluster. A removed file goes to .Trash, and fsck and dfsadmin -report show where each block lives and what each DataNode holds.

Experiment 3 — Hadoop's architecture, from its logs

1. Question

Demonstrate the Hadoop architecture components — HDFS, YARN and MapReduce — using sample logs.

2. Aim

Run one MapReduce job over a sample log, then follow it through the logs of the daemons that ran it.

3. Steps

On the cluster, 03_architecture.sh:

  1. List the daemons.
  2. Run a job, and find it.
  3. Read each daemon's log.

THE LOG TRACE TO FOLLOW

RM: application submitted → NM: AM container started → RM: map containers assigned → NM: map tasks start → NameNode: block reads served locally → RM: reduce containers assigned → NM: reduce fetches map outputs, over HTTP → RM: SUCCEEDED.

The one line worth finding is the shuffle fetch. It is the only step where data crosses the network in bulk, and it is what the combiner in experiment 7 exists to shrink.

4. Programme

On the cluster, 03_architecture.sh:

# Experiment 3 -- demonstrate the Hadoop architecture components using sample logs
#
# Run it: bash 03_architecture.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 06_yarn_scheduling.py, which runs the scheduler the logs describe
#
# Step 1: List the daemons
jps | awk '{print $2}' | sort         # the daemons, one JVM each

# Step 2: Run a job, and find it
hdfs dfs -mkdir -p /user/student/logs
hdfs dfs -put syslog /user/student/logs/
yarn jar $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar \
    wordcount /user/student/logs /user/student/wc-out 2>&1 | grep -E "Running job|completed successfully"
# [Corrected: this put /var/log/syslog, which many systems no longer have
# (journald keeps the log instead), into a directory that did not exist yet.
# Any text file will do; syslog here is a few lines of the lab's documents.]

APP=$(yarn application -list -appStates ALL 2>/dev/null | awk '/^application_/{print $1}' | tail -1)
yarn application -list -appStates ALL 2>/dev/null | awk -F'\t' 'NR>1{print $1 " | " $2 " | " $6 " | " $7}'
yarn application -status $APP 2>/dev/null | grep -E "Application-Name|State|Final-State"
sleep 10                              # the NodeManager uploads the logs when the job ends
yarn logs -applicationId $APP 2>/dev/null | grep "^Container: " | sort -u | sed -E 's/ on .*//'
# [Corrected: the commands used application_1699999999999_0001, an id from
# someone else's cluster. Every cluster numbers its own; take it from
# `yarn application -list`.]

# Step 3: Read each daemon's log
# [Corrected: these were `tail -f`, which never returns -- each would have to
# be stopped with Ctrl-C before the next. grep finds the lines described. And
# Hadoop 3 names the YARN logs hadoop-<user>-resourcemanager-<host>.log, not yarn-.]
grep -h -m 2 "BLOCK\* allocate" $HADOOP_LOG_DIR/hadoop-*-namenode-*.log | cut -c 25-
#   BlockStateChange lines: every allocation and replication decision
grep -h -m 1 "Receiving BP-" $HADOOP_LOG_DIR/hadoop-*-datanode-*.log | cut -c 25-
#   "Receiving BP-...:blk_..." -- a block landing, with its pipeline
grep -h -m 2 "Assigned container container_" $HADOOP_LOG_DIR/*-resourcemanager-*.log | cut -c 25-
#   "Assigned container container_..." -- the scheduler, deciding
grep -h -m 1 "Starting resource-monitoring for container_" $HADOOP_LOG_DIR/*-nodemanager-*.log | cut -c 25-
#   "Starting resource-monitoring for container_..." -- the container's life
yarn logs -applicationId $APP 2>/dev/null | grep -m 1 -oE "fetcher#[0-9]+ about to shuffle output of map [^ ]+"
#   in the reduce container's own log: one of the reduce's fetchers asks the
#   NodeManager's shuffle service, over HTTP, for a map's output -- THE SHUFFLE

# --- the trace to follow, in order -----------------------------------------
#   RM log         : application submitted, ApplicationMaster container assigned
#   NM log         : AM container started
#   RM log         : AM requests N map containers; scheduler assigns them
#   NM logs        : each map task starts, reports progress
#   NameNode log   : block reads served, LOCAL where possible
#   RM log         : reduce containers assigned after map progress passes 5%
#   NM log         : reduce fetches map outputs -- THE SHUFFLE, over HTTP
#   RM log         : application FINISHED, SUCCEEDED
#
# The one line worth finding is the shuffle fetch. It is the only step where
# data crosses the network in bulk, and it is what the combiner in
# experiment 7 exists to shrink.

5. Execution and Results

On the cluster, 03_architecture.sh:

OUTPUT

$ jps | awk '{print $2}' | sort
DataNode
DataNode
DataNode
DataNode
JobHistoryServer
Jps
NameNode
NodeManager
ResourceManager
SecondaryNameNode
$ hdfs dfs -mkdir -p /user/student/logs
$ hdfs dfs -put syslog /user/student/logs/
$ yarn jar $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar \
      wordcount /user/student/logs /user/student/wc-out 2>&1 | grep -E "Running job|completed successfully"
2026-10-04 23:40:14,246 INFO mapreduce.Job: Running job: job_1791157202192_0001
2026-10-04 23:40:31,543 INFO mapreduce.Job: Job job_1791157202192_0001 completed successfully
$ APP=$(yarn application -list -appStates ALL 2>/dev/null | awk '/^application_/{print $1}' | tail -1)
$ yarn application -list -appStates ALL 2>/dev/null | awk -F'\t' 'NR>1{print $1 " | " $2 " | " $6 " | " $7}'
                Application-Id |     Application-Name |              State |        Final-State
application_1791157202192_0001 |           word count |           FINISHED |          SUCCEEDED
$ yarn application -status $APP 2>/dev/null | grep -E "Application-Name|State|Final-State"
    Application-Name : word count
    State : FINISHED
    Final-State : SUCCEEDED
$ sleep 10
$ yarn logs -applicationId $APP 2>/dev/null | grep "^Container: " | sort -u | sed -E 's/ on .*//'
Container: container_1791157202192_0001_01_000001
Container: container_1791157202192_0001_01_000002
Container: container_1791157202192_0001_01_000003
$ grep -h -m 2 "BLOCK\* allocate" $HADOOP_LOG_DIR/hadoop-*-namenode-*.log | cut -c 25-
INFO org.apache.hadoop.hdfs.StateChange: BLOCK* allocate blk_1073741825_1001, replicas=127.0.0.1:9886, 127.0.0.1:9866, 127.0.0.1:9896 for /user/student/logs/syslog._COPYING_
INFO org.apache.hadoop.hdfs.StateChange: BLOCK* allocate blk_1073741826_1002, replicas=127.0.0.1:9886, 127.0.0.1:9896, 127.0.0.1:9906 for /tmp/hadoop-yarn/staging/root/.staging/job_1791157202192_0001/job.jar
$ grep -h -m 1 "Receiving BP-" $HADOOP_LOG_DIR/hadoop-*-datanode-*.log | cut -c 25-
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:57158 dest: /127.0.0.1:9886
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:35392 dest: /127.0.0.1:9896
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741826_1002 src: /127.0.0.1:32904 dest: /127.0.0.1:9906
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:42626 dest: /127.0.0.1:9866
$ grep -h -m 2 "Assigned container container_" $HADOOP_LOG_DIR/*-resourcemanager-*.log | cut -c 25-
INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.common.fica.FiCaSchedulerNode: Assigned container container_1791157202192_0001_01_000001 of capacity <memory:512, vCores:1> on host localhost:40995, which has 1 containers, <memory:512, vCores:1> used and <memory:5632, vCores:3> available after allocation
INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.common.fica.FiCaSchedulerNode: Assigned container container_1791157202192_0001_01_000002 of capacity <memory:512, vCores:1> on host localhost:40995, which has 2 containers, <memory:1024, vCores:2> used and <memory:5120, vCores:2> available after allocation
$ grep -h -m 1 "Starting resource-monitoring for container_" $HADOOP_LOG_DIR/*-nodemanager-*.log | cut -c 25-
INFO org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl: Starting resource-monitoring for container_1791157202192_0001_01_000001
$ yarn logs -applicationId $APP 2>/dev/null | grep -m 1 -oE "fetcher#[0-9]+ about to shuffle output of map [^ ]+"
fetcher#2 about to shuffle output of map attempt_1791157202192_0001_m_000000_0

RESULT

One job, traced through four daemons: the ResourceManager admitted it and allocated its containers, the NodeManager launched them, the NameNode served the blocks, and the reduce fetched the map outputs over HTTP — the shuffle.

Experiment 4 — Storing a large file in HDFS: blocks and replication

1. Question

Store and retrieve a large file in HDFS, and demonstrate its block distribution and replication factor.

2. Aim

Put a 300 MB file into HDFS, find its blocks and where each replica is, change its replication, and work out the block arithmetic and the small-files cost.

3. Steps

On the cluster, 04_hdfs_store.sh:

  1. Make a 300 MB file.
  2. Put it in HDFS.
  3. Find its blocks and their replicas.
  4. Change its replication.
  5. Write it again with 64 MB blocks.
  6. Get it back, and compare.

The Python check, 04_blocks_replication.py:

  1. Size the blocks.
  2. Place a 1 GB file's replicas.
  3. Count the small-files cost.

THE BLOCK TABLE

File Blocks Last block Disk used
1 MB 1 1.00 MB 1 MB
128 MB 1 128.00 MB 128 MB
129 MB 2 1.00 MB 129 MB
260 MB 3 4.00 MB 260 MB
1024 MB 8 128.00 MB 1024 MB
5000 MB 40 8.00 MB 5000 MB

A 260 MB file is 128 + 128 + 4. HDFS wastes no space on block padding, and this is the most examined calculation in the course.

But 128 MB → 1 block and 129 MB → 2. One byte past the boundary costs a whole block object in NameNode RAM, though almost no disk.

4. Programme

On the cluster, 04_hdfs_store.sh:

# Experiment 4 -- store and retrieve large files in HDFS -- block distribution and replication
#
# Run it: bash 04_hdfs_store.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 04_blocks_replication.py, which computes every figure below
#
# make a file bigger than one block so there is something to distribute
# Step 1: Make a 300 MB file
dd if=/dev/urandom of=big.bin bs=1M count=300      # 300 MB -> 3 blocks

# Step 2: Put it in HDFS
hdfs dfs -mkdir -p /user/student/big
hdfs dfs -put big.bin /user/student/big/

# how many blocks, and where are they?
# Step 3: Find its blocks and their replicas
hdfs fsck /user/student/big/big.bin -files -blocks -locations
#   expect: 3 blocks -- 128 MB, 128 MB, 44 MB
#   the LAST BLOCK IS SHORT. HDFS does not pad.

# Step 4: Change its replication
hdfs dfs -stat "%r" /user/student/big/big.bin      # replication factor
hdfs dfs -setrep -w 2 /user/student/big/big.bin    # -w waits for completion
hdfs fsck /user/student/big/big.bin -files -blocks # now 2 locations per block

# a non-default block size, set PER FILE at write time
# Step 5: Write it again with 64 MB blocks
hdfs dfs -D dfs.blocksize=67108864 -put big.bin /user/student/big/small-blocks.bin
hdfs fsck /user/student/big/small-blocks.bin -files -blocks
#   expect: 5 blocks -- four of 64 MB and a last one of 44 MB -- more blocks,
#   more NameNode objects, more map tasks (one per block by default)
# [Corrected: this said "5 blocks of 64 MB". 300 MB is 4 x 64 + 44: the last
# block is short here too, as it is above.]

# retrieve and verify
# Step 6: Get it back, and compare
hdfs dfs -get /user/student/big/big.bin ./back.bin
md5sum big.bin back.bin                            # must match

hdfs dfsadmin -report | grep -E "Name|DFS Used|Remaining"

The Python check, 04_blocks_replication.py:

"""Experiment 4 -- store and retrieve a large file in HDFS: blocks, block
distribution and the replication factor.

`04_hdfs_store.sh` carries the commands you actually type, and runs on a real
cluster (the lab page shows it). What runs here is the ARITHMETIC, which is the
part that gets examined and the part students get wrong.
"""
from blocks import BLOCK, MB, blocks_for, namenode_memory, placement


def main():
    print("  Experiment 4 -- HDFS blocks, distribution and replication")

    # Step 1: Size the blocks
    print("\n    block sizing (default block = 128 MB):")
    print(f"    {'file':>10}  {'blocks':>6}  {'last block':>12}  {'disk used':>12}")
    for mb in (1, 128, 129, 260, 1024, 5000):
        n, last = blocks_for(mb * MB)
        print(f"    {mb:>7} MB  {n:>6}  {last / MB:>9.2f} MB  {mb:>9} MB")
    print("""         a 260 MB file is 128 + 128 + 4, NOT three full blocks.
         An HDFS block is a logical MAXIMUM; the last block occupies
         only what it needs. HDFS wastes no space on block padding --
         which is the opposite of what the word 'block' suggests, and
         the mistake to avoid in the exam""")

    n1, last1 = blocks_for(1 * MB)
    assert (n1, last1) == (1, 1 * MB)
    n260, last260 = blocks_for(260 * MB)
    assert n260 == 3 and last260 == 4 * MB
    n128, _ = blocks_for(128 * MB)
    n129, _ = blocks_for(129 * MB)
    assert (n128, n129) == (1, 2), "one byte over a block boundary costs a block"
    print("\n    128 MB -> 1 block, 129 MB -> 2 blocks")
    print("""         one byte past the boundary costs a whole block OBJECT
         in NameNode memory, though almost no disk. Metadata is the
         scarce resource in HDFS, not disk""")

    # Step 2: Place a 1 GB file's replicas
    print("\n    a 1 GB file, replication 3, 6 DataNodes across 2 racks:")
    n, last = blocks_for(1024 * MB)
    plan, rack_of = placement(n, 3, datanodes=6, racks=2)
    print(f"    {n} blocks (last = {last / MB:.0f} MB)")
    print(f"    {'block':>6}  {'replicas (node/rack)':<34}  racks used")
    for i, nodes in enumerate(plan):
        desc = "  ".join(f"n{d}/r{rack_of[d]}" for d in nodes)
        print(f"    {i:>6}  {desc:<34}  {len({rack_of[d] for d in nodes})}")
    for nodes in plan:
        assert len({rack_of[d] for d in nodes}) == 2, "must span two racks"
        assert len(set(nodes)) == 3, "three replicas on three distinct nodes"
    print("""         every block spans EXACTLY TWO racks: one replica on the
         writer's rack, two on another. Two racks survive a rack
         failure; a third rack would double cross-rack write traffic
         to buy very little. That trade is the whole policy""")

    raw = 1024
    print(f"\n    storage cost: {raw} MB of data at replication 3 "
          f"occupies {raw * 3} MB of disk")
    print(f"    the same data with erasure coding (RS-6-3) would occupy "
          f"{raw * 9 // 6} MB")
    assert raw * 3 == 3072 and raw * 9 // 6 == 1536
    print("""         replication costs 200% overhead for 3x durability;
         RS-6-3 erasure coding costs 50% for comparable durability,
         at the price of expensive reconstruction reads. HDFS added
         erasure coding in 3.0 for exactly this reason -- COLD data""")

    # Step 3: Count the small-files cost
    print("\n    the small-files problem, in NameNode RAM (~150 bytes/object):")
    print(f"    {'scenario':<26}{'files':>12}{'blocks':>12}{'NameNode RAM':>16}")
    for label, files, size_each in (
            ("one 1 GB file", 1, 1024 * MB),
            ("1,000 x 1 MB files", 1000, 1 * MB),
            ("1,000,000 x 1 KB files", 1_000_000, 1024)):
        total_blocks = sum(blocks_for(size_each)[0] for _ in range(1)) * files
        ram = namenode_memory(files, total_blocks)
        print(f"    {label:<26}{files:>12,}{total_blocks:>12,}"
              f"{ram / MB:>13.2f} MB")
    one = namenode_memory(1, 8)
    many = namenode_memory(1_000_000, 1_000_000)
    assert many // one > 200_000
    print(f"""         the same 1 GB costs {one} bytes as one file and
         {many / MB:.0f} MB as a million small ones -- a factor of
         {many // one:,}. HDFS was built for few large files, and this
         single table is the reason""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 04_hdfs_store.sh:

OUTPUT

$ dd if=/dev/urandom of=big.bin bs=1M count=300
300+0 records in
300+0 records out
314572800 bytes (315 MB, 300 MiB) copied, 2.09212 s, 150 MB/s
$ hdfs dfs -mkdir -p /user/student/big
$ hdfs dfs -put big.bin /user/student/big/
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&locations=1&path=%2Fuser%2Fstudent%2Fbig%2Fbig.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/big.bin at Sun Oct 04 23:08:33 UTC 2026

/user/student/big/big.bin 314572800 bytes, replicated: replication=3, 3 block(s):  OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741825_1001 len=134217728 Live_repl=3  [DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-1431a54e-afc3-4c6b-bc5f-702712a1ce03,DISK]]
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741826_1002 len=134217728 Live_repl=3  [DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-1431a54e-afc3-4c6b-bc5f-702712a1ce03,DISK]]
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741827_1003 len=46137344 Live_repl=3  [DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9886,DS-2d1746a3-8235-40de-befd-514261a60ca2,DISK]]


Status: HEALTHY
 Number of data-nodes:  4
 Number of racks:       1
 Total dirs:            0
 Total symlinks:        0

Replicated Blocks:
 Total size:    314572800 B
 Total files:   1
 Total blocks (validated):  3 (avg. block size 104857600 B)
 Minimally replicated blocks:   3 (100.0 %)
 Over-replicated blocks:    0 (0.0 %)
 Under-replicated blocks:   0 (0.0 %)
 Mis-replicated blocks:     0 (0.0 %)
 Default replication factor:    3
 Average block replication: 3.0
 Missing blocks:        0
 Corrupt blocks:        0
 Missing replicas:      0 (0.0 %)
 Blocks queued for replication: 0

Erasure Coded Block Groups:
 Total size:    0 B
 Total files:   0
 Total block groups (validated):    0
 Minimally erasure-coded block groups:  0
 Over-erasure-coded block groups:   0
 Under-erasure-coded block groups:  0
 Unsatisfactory placement block groups: 0
 Average block group size:  0.0
 Missing block groups:      0
 Corrupt block groups:      0
 Missing internal blocks:   0
 Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:33 UTC 2026 in 10 milliseconds


The filesystem under path '/user/student/big/big.bin' is HEALTHY
$ hdfs dfs -stat "%r" /user/student/big/big.bin
3
$ hdfs dfs -setrep -w 2 /user/student/big/big.bin
Replication 2 set: /user/student/big/big.bin
Waiting for /user/student/big/big.bin ...
WARNING: the waiting time may be long for DECREASING the number of replications.
. done
$ hdfs fsck /user/student/big/big.bin -files -blocks
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&path=%2Fuser%2Fstudent%2Fbig%2Fbig.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/big.bin at Sun Oct 04 23:08:48 UTC 2026

/user/student/big/big.bin 314572800 bytes, replicated: replication=2, 3 block(s):  OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741825_1001 len=134217728 Live_repl=2
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741826_1002 len=134217728 Live_repl=2
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741827_1003 len=46137344 Live_repl=2


Status: HEALTHY
 Number of data-nodes:  4
 Number of racks:       1
 Total dirs:            0
 Total symlinks:        0

Replicated Blocks:
 Total size:    314572800 B
 Total files:   1
 Total blocks (validated):  3 (avg. block size 104857600 B)
 Minimally replicated blocks:   3 (100.0 %)
 Over-replicated blocks:    0 (0.0 %)
 Under-replicated blocks:   0 (0.0 %)
 Mis-replicated blocks:     0 (0.0 %)
 Default replication factor:    3
 Average block replication: 2.0
 Missing blocks:        0
 Corrupt blocks:        0
 Missing replicas:      0 (0.0 %)
 Blocks queued for replication: 0

Erasure Coded Block Groups:
 Total size:    0 B
 Total files:   0
 Total block groups (validated):    0
 Minimally erasure-coded block groups:  0
 Over-erasure-coded block groups:   0
 Under-erasure-coded block groups:  0
 Unsatisfactory placement block groups: 0
 Average block group size:  0.0
 Missing block groups:      0
 Corrupt block groups:      0
 Missing internal blocks:   0
 Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:48 UTC 2026 in 1 milliseconds


The filesystem under path '/user/student/big/big.bin' is HEALTHY
$ hdfs dfs -D dfs.blocksize=67108864 -put big.bin /user/student/big/small-blocks.bin
$ hdfs fsck /user/student/big/small-blocks.bin -files -blocks
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&path=%2Fuser%2Fstudent%2Fbig%2Fsmall-blocks.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/small-blocks.bin at Sun Oct 04 23:08:55 UTC 2026

/user/student/big/small-blocks.bin 314572800 bytes, replicated: replication=3, 5 block(s):  OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741828_1004 len=67108864 Live_repl=3
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741829_1005 len=67108864 Live_repl=3
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741830_1006 len=67108864 Live_repl=3
3. BP-245208533-127.0.0.1-1791155276951:blk_1073741831_1007 len=67108864 Live_repl=3
4. BP-245208533-127.0.0.1-1791155276951:blk_1073741832_1008 len=46137344 Live_repl=3


Status: HEALTHY
 Number of data-nodes:  4
 Number of racks:       1
 Total dirs:            0
 Total symlinks:        0

Replicated Blocks:
 Total size:    314572800 B
 Total files:   1
 Total blocks (validated):  5 (avg. block size 62914560 B)
 Minimally replicated blocks:   5 (100.0 %)
 Over-replicated blocks:    0 (0.0 %)
 Under-replicated blocks:   0 (0.0 %)
 Mis-replicated blocks:     0 (0.0 %)
 Default replication factor:    3
 Average block replication: 3.0
 Missing blocks:        0
 Corrupt blocks:        0
 Missing replicas:      0 (0.0 %)
 Blocks queued for replication: 0

Erasure Coded Block Groups:
 Total size:    0 B
 Total files:   0
 Total block groups (validated):    0
 Minimally erasure-coded block groups:  0
 Over-erasure-coded block groups:   0
 Under-erasure-coded block groups:  0
 Unsatisfactory placement block groups: 0
 Average block group size:  0.0
 Missing block groups:      0
 Corrupt block groups:      0
 Missing internal blocks:   0
 Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:55 UTC 2026 in 2 milliseconds


The filesystem under path '/user/student/big/small-blocks.bin' is HEALTHY
$ hdfs dfs -get /user/student/big/big.bin ./back.bin
$ md5sum big.bin back.bin
6792d150c0ab35807e5020be57771bb5  big.bin
6792d150c0ab35807e5020be57771bb5  back.bin
$ hdfs dfsadmin -report | grep -E "Name|DFS Used|Remaining"
DFS Remaining: 29910220800 (27.86 GB)
DFS Used: 1585250451 (1.48 GB)
DFS Used%: 5.03%
Name: 127.0.0.1:9866 (localhost)
DFS Used: 498819114 (475.71 MB)
Non DFS Used: 31804215254 (29.62 GB)
DFS Remaining: 7477854208 (6.96 GB)
DFS Used%: 0.18%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9886 (localhost)
DFS Used: 181788693 (173.37 MB)
Non DFS Used: 32121655275 (29.92 GB)
DFS Remaining: 7477444608 (6.96 GB)
DFS Used%: 0.07%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9896 (localhost)
DFS Used: 587587633 (560.37 MB)
Non DFS Used: 31715823567 (29.54 GB)
DFS Remaining: 7477477376 (6.96 GB)
DFS Used%: 0.22%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9906 (localhost)
DFS Used: 317055011 (302.37 MB)
Non DFS Used: 31986388957 (29.79 GB)
DFS Remaining: 7477444608 (6.96 GB)
DFS Used%: 0.12%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%

The Python check, 04_blocks_replication.py:

OUTPUT

  Experiment 4 -- HDFS blocks, distribution and replication

    block sizing (default block = 128 MB):
          file  blocks    last block     disk used
          1 MB       1       1.00 MB          1 MB
        128 MB       1     128.00 MB        128 MB
        129 MB       2       1.00 MB        129 MB
        260 MB       3       4.00 MB        260 MB
       1024 MB       8     128.00 MB       1024 MB
       5000 MB      40       8.00 MB       5000 MB
         a 260 MB file is 128 + 128 + 4, NOT three full blocks.
         An HDFS block is a logical MAXIMUM; the last block occupies
         only what it needs. HDFS wastes no space on block padding --
         which is the opposite of what the word 'block' suggests, and
         the mistake to avoid in the exam

    128 MB -> 1 block, 129 MB -> 2 blocks
         one byte past the boundary costs a whole block OBJECT
         in NameNode memory, though almost no disk. Metadata is the
         scarce resource in HDFS, not disk

    a 1 GB file, replication 3, 6 DataNodes across 2 racks:
    8 blocks (last = 128 MB)
     block  replicas (node/rack)                racks used
         0  n0/r0  n1/r1  n3/r1                 2
         1  n2/r0  n3/r1  n5/r1                 2
         2  n4/r0  n5/r1  n1/r1                 2
         3  n0/r0  n1/r1  n5/r1                 2
         4  n2/r0  n3/r1  n1/r1                 2
         5  n4/r0  n5/r1  n3/r1                 2
         6  n0/r0  n1/r1  n3/r1                 2
         7  n2/r0  n3/r1  n5/r1                 2
         every block spans EXACTLY TWO racks: one replica on the
         writer's rack, two on another. Two racks survive a rack
         failure; a third rack would double cross-rack write traffic
         to buy very little. That trade is the whole policy

    storage cost: 1024 MB of data at replication 3 occupies 3072 MB of disk
    the same data with erasure coding (RS-6-3) would occupy 1536 MB
         replication costs 200% overhead for 3x durability;
         RS-6-3 erasure coding costs 50% for comparable durability,
         at the price of expensive reconstruction reads. HDFS added
         erasure coding in 3.0 for exactly this reason -- COLD data

    the small-files problem, in NameNode RAM (~150 bytes/object):
    scenario                         files      blocks    NameNode RAM
    one 1 GB file                        1           8         0.00 MB
    1,000 x 1 MB files               1,000       1,000         0.29 MB
    1,000,000 x 1 KB files       1,000,000   1,000,000       286.10 MB
         the same 1 GB costs 1350 bytes as one file and
         286 MB as a million small ones -- a factor of
         222,222. HDFS was built for few large files, and this
         single table is the reason

REPLICA PLACEMENT, 1 GB OVER 6 NODES IN 2 RACKS

Block Replicas Racks
0 n0/r0, n1/r1, n3/r1 2
1 n2/r0, n3/r1, n5/r1 2
2 n4/r0, n5/r1, n1/r1 2
… … 2

Every block spans exactly two racks — one replica on the writer's rack, two on another. Asserted for all 8 blocks.

THE SMALL-FILES TABLE

Scenario Files Blocks NameNode RAM
one 1 GB file 1 8 0.00 MB (1,350 bytes)
1,000 × 1 MB 1,000 1,000 0.29 MB
1,000,000 × 1 KB 1,000,000 1,000,000 286.10 MB

A factor of 222,222 for the same gigabyte.

AND THE STORAGE TRADE

1 GB at replication 3 occupies 3,072 MB; under RS-6-3 erasure coding, 1,536 MB — 200% overhead against 50%, at the cost of expensive reconstruction reads.

RESULT

A 300 MB file is 128 + 128 + 44 MB, three blocks, each on three of the four DataNodes; the same file in 64 MB blocks is five. The arithmetic agrees: 260 MB is 128 + 128 + 4, and a million 1 KB files cost the NameNode 222,222 times the memory of one 1 GB file.

Experiment 5 — Fault tolerance and recovery

1. Question

Simulate NameNode and DataNode failure, and observe fault tolerance and recovery.

2. Aim

Kill a DataNode and read the file anyway; wait for the NameNode to declare it dead and re-replicate; kill the NameNode and watch it recover from its image and edits; then work out which failures lose data.

3. Steps

On the cluster, 05_fault_tolerance.sh:

  1. Have the 300 MB file.
  2. Kill a DataNode, and read the file.
  3. Wait for it to be declared dead.
  4. Bring it back.
  5. Kill the NameNode, and restart it.
  6. Read what recovery reads.

The Python check, 05_fault_tolerance.py:

  1. Fail nodes and racks.
  2. Find the worst case.
  3. Re-replicate.
  4. Lose the NameNode.
  5. Compare the three answers.

WHICH FAILURES LOSE DATA

Failure Blocks live Blocks lost
1 DataNode (n1) 8 0
2 DataNodes (n1, n3) 8 0
3 DataNodes (n1, n3, n5) 8 0
3 DataNodes (n0, n1, n3) 6 2
a whole rack (r1) 8 0
both racks 0 8

Rows 3 and 4 are the point. n1, n3, n5 are rack 1, so rows 3 and 5 are the same failure written two ways — and both are survivable. Three failures only hurt when they straddle the racks.

4. Programme

On the cluster, 05_fault_tolerance.sh:

# Experiment 5 -- simulate NameNode/DataNode failure and observe fault tolerance and recovery
# Run it: bash 05_fault_tolerance.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
# The runnable half is 05_fault_tolerance.py, which models which blocks survive which failures
# Step 1: Have the 300 MB file
hdfs dfs -test -e /user/student/big/big.bin || {
  dd if=/dev/urandom of=big.bin bs=1M count=300 status=none
  hdfs dfs -mkdir -p /user/student/big && hdfs dfs -put big.bin /user/student/big/; }

# Step 2: Kill a DataNode, and read the file
hdfs dfsadmin -report 2>/dev/null | grep -E "^Name:|Live datanodes"
hdfs fsck /user/student/big/big.bin -files -blocks -locations > before.txt 2>/dev/null
grep -o "DatanodeInfoWithStorage" before.txt | wc -l   # 3 blocks x 3 replicas = 9

# kill one DataNode. jps shows four here -- this cluster runs four on one
# machine, as a multi-node cluster runs one per host -- so pick it by its pid file
jps | grep -c DataNode
kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-datanode.pid)
# [Corrected: `kill -9 <datanode_pid>` was a placeholder. A DataNode writes its
# pid to $HADOOP_PID_DIR (by default /tmp) as hadoop-<user>-datanode.pid.]

# the file is STILL READABLE, immediately -- other replicas serve it
hdfs dfs -cat /user/student/big/big.bin 2>client.log | wc -c
#   every byte. A client tries a block's replicas in turn: if it tries the dead
#   node first it logs a warning (kept in client.log here) and reads another
#   replica. Which it tries first differs from read to read.

# the NameNode does not react at once: dfs.heartbeat.interval (3 s) and
# dfs.namenode.heartbeat.recheck-interval (5 min) give
#   10 * 3 + 2 * 300 = 630 seconds  before the node is declared DEAD
# This cluster sets the recheck interval to 15 s, so 10 * 3 + 2 * 15 = 60 s:
# Step 3: Wait for it to be declared dead
sleep 75
hdfs dfsadmin -report -dead 2>/dev/null | grep -E "Dead datanodes"
hdfs fsck / 2>/dev/null | grep -E "Under-replicated|Missing blocks:"
#   under-replicated blocks appear, then disappear as HDFS re-replicates --
#   on a cluster this small, within seconds of the node being declared dead,
#   so by now the copies are made. fsck has where every replica now lives:
hdfs fsck /user/student/big/big.bin -files -blocks -locations 2>/dev/null \
  | grep -oE "DatanodeInfoWithStorage\[[0-9.:]+" | sort | uniq -c
#   3 blocks x 3 replicas on the three nodes still alive: each has all three.
#   Whichever blocks the dead node held were copied again.
# [Corrected: this counted the NameNode's "to replicate blk_" log lines, as
# Hadoop 2 logged each order. Hadoop 3 logs them only at DEBUG, so the count
# is 0 however many blocks were copied; where the replicas are is the evidence.]
# [Corrected: this waited 660 s, for the default 630 s. The wait follows the
# setting; dfsadmin -report -dead lists only the dead.]

# Step 4: Bring it back
# bring it back
$HADOOP_HOME/bin/hdfs --daemon start datanode
sleep 15
hdfs fsck / 2>/dev/null | grep -E "Over-replicated"    # briefly over-replicated, then trimmed

# Step 5: Kill the NameNode, and restart it
jps | grep -c NameNode              # NameNode and SecondaryNameNode
kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-namenode.pid)
hdfs dfs -ls / 2>&1 | grep -m 1 -o "Call From .* failed on connection exception"
                            # FAILS. The cluster is unusable. Nothing was lost,
                            # but nothing is reachable either.

$HADOOP_HOME/bin/hdfs --daemon start namenode
sleep 5
hdfs dfsadmin -safemode get # ON -- it is collecting block reports
hdfs dfsadmin -safemode wait
#   safe mode leaves once dfs.namenode.safemode.threshold-pct (0.999) of
#   blocks have reported. On a large cluster this takes MINUTES, and it is
#   why HA exists.
hdfs dfs -cat /user/student/big/big.bin | wc -c   # every byte, still

# Step 6: Read what recovery reads
ls $(hdfs getconf -confKey dfs.namenode.name.dir | sed 's|^file://||')/current/ | sed -E 's/[0-9]{19}/N/g' | sort -u
#   fsimage_N                      the namespace at a checkpoint
#   edits_N-N, edits_inprogress_N  every change since
#   VERSION                        clusterID -- must match the DataNodes'
# [Corrected: this listed /usr/local/hadoop_store/hdfs/namenode/current, the
# directory experiment 1 configured on its machine. hdfs getconf asks the
# cluster where its own is; the transaction numbers are shown as N.]
# The BLOCK MAP is in NONE of these. It is rebuilt from block reports.

The Python check, 05_fault_tolerance.py:

"""Experiment 5 -- simulate NameNode/DataNode failure and observe fault
tolerance and recovery.

`05_fault_tolerance.sh` carries the commands. What runs here is the model:
which blocks survive which failures, and why the NameNode is the one failure
that is different in kind.
"""
import itertools

from blocks import blocks_for, placement


def surviving(plan, rack_of, dead_nodes=(), dead_racks=()):
    """Blocks with at least one live replica, and blocks fully lost."""
    dead = set(dead_nodes) | {d for d in rack_of if rack_of[d] in dead_racks}
    live, lost = [], []
    for i, nodes in enumerate(plan):
        (live if any(n not in dead for n in nodes) else lost).append(i)
    return live, lost


def main():
    print("  Experiment 5 -- fault tolerance and recovery")

    # Step 1: Fail nodes and racks
    n, _ = blocks_for(1024 * 1024 * 1024)
    plan, rack_of = placement(n, 3, datanodes=6, racks=2)
    print(f"\n    a 1 GB file: {n} blocks, replication 3, 6 nodes, 2 racks")

    print(f"\n    {'failure':<28}{'blocks live':>12}{'blocks lost':>12}  verdict")
    scenarios = [
        ("1 DataNode  (n1)",        dict(dead_nodes=[1])),
        ("2 DataNodes (n1, n3)",    dict(dead_nodes=[1, 3])),
        ("3 DataNodes (n1, n3, n5)", dict(dead_nodes=[1, 3, 5])),
        ("3 DataNodes (n0, n1, n3)", dict(dead_nodes=[0, 1, 3])),
        ("a whole rack (r1)",       dict(dead_racks=[1])),
        ("both racks",              dict(dead_racks=[0, 1])),
    ]
    for label, kw in scenarios:
        live, lost = surviving(plan, rack_of, **kw)
        verdict = "no data loss" if not lost else f"DATA LOSS on {len(lost)}"
        print(f"    {label:<28}{len(live):>12}{len(lost):>12}  {verdict}")

    live, lost = surviving(plan, rack_of, dead_racks=[1])
    assert lost == [], "losing one whole rack must not lose data"
    live, lost = surviving(plan, rack_of, dead_nodes=[1, 3, 5])
    assert lost == [], "n1, n3, n5 IS rack 1 -- the same failure, renamed"
    print("""         losing an ENTIRE RACK loses nothing, because every block
         keeps one replica on the other rack. That is precisely what
         the placement policy bought, and it is the answer to 'why
         rack awareness?'
         Note rows 3 and 4: n1, n3, n5 ARE rack 1, so those are the
         same failure written two ways -- and both are survivable.
         Three failures only hurt when they straddle the racks, as
         (n0, n1, n3) does""")

    # Step 2: Find the worst case
    print("\n    the worst case, by brute force -- how many DataNode failures")
    print("    can this layout survive with certainty?")
    worst = None
    for k in range(1, 7):
        bad = [combo for combo in itertools.combinations(range(6), k)
               if surviving(plan, rack_of, dead_nodes=combo)[1]]
        total = len(list(itertools.combinations(range(6), k)))
        print(f"      {k} node(s) down: {len(bad):>3} of {total:>3} "
              f"combinations lose data")
        if bad and worst is None:
            worst = k
    assert worst == 3, "replication 3 tolerates ANY 2 failures, not any 3"
    print("""         ANY TWO failures are survivable; some threes are not.
         Replication factor R tolerates R-1 arbitrary failures --
         and note that most 3-node combinations are still fine, so
         'replication 3 fails at 3 nodes' is only true of the worst
         case, which is the honest way to state it""")

    # Step 3: Re-replicate
    print("\n    re-replication after a DataNode is declared dead:")
    print("      1. DataNode misses heartbeats (default: 3 sec interval)")
    print("      2. NameNode waits 10 * 3 sec + 2 * 5 min = 10 min 30 sec")
    print("      3. its blocks are now UNDER-REPLICATED (2 of 3)")
    print("      4. NameNode schedules copies from surviving replicas")
    print("      5. replication returns to 3; no client ever saw an error")
    stale = 10 * 3 + 2 * 5 * 60
    assert stale == 630
    print(f"""         the {stale}-second default is deliberately LONG. A node
         that reboots in five minutes should not trigger a cluster-wide
         copy storm, so HDFS trades a longer window of reduced
         redundancy for far less needless network traffic""")

    # Step 4: Lose the NameNode
    print("\n    the NameNode is a different kind of failure:")
    print(f"      {'component':<22}{'holds':<34}{'lost on crash?'}")
    for comp, holds, lost_ in (
            ("fsimage (on disk)", "the namespace at a checkpoint", "no"),
            ("edit log (on disk)", "changes since the checkpoint", "no"),
            ("block map (in RAM)", "block -> DataNode locations", "YES"),
    ):
        print(f"      {comp:<22}{holds:<34}{lost_}")
    print("""         the block MAP is never persisted -- it is rebuilt from
         DataNode block reports at startup, which is why a large
         NameNode takes minutes to leave safe mode. The namespace
         survives; the locations are reconstructed""")

    # Step 5: Compare the three answers
    print("\n    the three answers to NameNode failure, in historical order:")
    print(f"      {'mechanism':<26}{'recovers':<16}{'automatic?'}")
    for m, r, a in (("Secondary NameNode", "checkpoint only", "no -- NOT a standby"),
                    ("NameNode HA (2 NNs)", "full", "yes, via ZooKeeper"),
                    ("HDFS Federation", "n/a -- scales namespace", "n/a")):
        print(f"      {m:<26}{r:<16}{a}")
    print("""         the Secondary NameNode is the most misleadingly named
         component in Hadoop: it merges fsimage with the edit log so
         restarts stay fast, and it CANNOT take over. HA needs two
         NameNodes, a shared edit log (QJM) and ZooKeeper for failover
         -- which is exactly why experiment 16 exists""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 05_fault_tolerance.sh:

OUTPUT

$ hdfs dfs -test -e /user/student/big/big.bin || {
    dd if=/dev/urandom of=big.bin bs=1M count=300 status=none
    hdfs dfs -mkdir -p /user/student/big && hdfs dfs -put big.bin /user/student/big/; }
$ hdfs dfsadmin -report 2>/dev/null | grep -E "^Name:|Live datanodes"
Live datanodes (4):
Name: 127.0.0.1:9866 (localhost)
Name: 127.0.0.1:9886 (localhost)
Name: 127.0.0.1:9896 (localhost)
Name: 127.0.0.1:9906 (localhost)
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations > before.txt 2>/dev/null
$ grep -o "DatanodeInfoWithStorage" before.txt | wc -l
9
$ jps | grep -c DataNode
4
$ kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-datanode.pid)
$ hdfs dfs -cat /user/student/big/big.bin 2>client.log | wc -c
314572800
$ sleep 75
$ hdfs dfsadmin -report -dead 2>/dev/null | grep -E "Dead datanodes"
Dead datanodes (1):
$ hdfs fsck / 2>/dev/null | grep -E "Under-replicated|Missing blocks:"
 Under-replicated blocks:   0 (0.0 %)
 Missing blocks:        0
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations 2>/dev/null \
    | grep -oE "DatanodeInfoWithStorage\[[0-9.:]+" | sort | uniq -c
      3 DatanodeInfoWithStorage[127.0.0.1:9886
      3 DatanodeInfoWithStorage[127.0.0.1:9896
      3 DatanodeInfoWithStorage[127.0.0.1:9906
$ $HADOOP_HOME/bin/hdfs --daemon start datanode
$ sleep 15
$ hdfs fsck / 2>/dev/null | grep -E "Over-replicated"
 Over-replicated blocks:    0 (0.0 %)
$ jps | grep -c NameNode
2
$ kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-namenode.pid)
$ hdfs dfs -ls / 2>&1 | grep -m 1 -o "Call From .* failed on connection exception"
Call From vm/127.0.0.1 to localhost:9000 failed on connection exception
$ $HADOOP_HOME/bin/hdfs --daemon start namenode
$ sleep 5
$ hdfs dfsadmin -safemode get
Safe mode is ON
$ hdfs dfsadmin -safemode wait
Safe mode is OFF
$ hdfs dfs -cat /user/student/big/big.bin | wc -c
314572800
$ ls $(hdfs getconf -confKey dfs.namenode.name.dir | sed 's|^file://||')/current/ | sed -E 's/[0-9]{19}/N/g' | sort -u
VERSION
edits_N-N
edits_inprogress_N
fsimage_N
fsimage_N.md5
seen_txid

The Python check, 05_fault_tolerance.py:

OUTPUT

  Experiment 5 -- fault tolerance and recovery

    a 1 GB file: 8 blocks, replication 3, 6 nodes, 2 racks

    failure                      blocks live blocks lost  verdict
    1 DataNode  (n1)                       8           0  no data loss
    2 DataNodes (n1, n3)                   8           0  no data loss
    3 DataNodes (n1, n3, n5)               8           0  no data loss
    3 DataNodes (n0, n1, n3)               6           2  DATA LOSS on 2
    a whole rack (r1)                      8           0  no data loss
    both racks                             0           8  DATA LOSS on 8
         losing an ENTIRE RACK loses nothing, because every block
         keeps one replica on the other rack. That is precisely what
         the placement policy bought, and it is the answer to 'why
         rack awareness?'
         Note rows 3 and 4: n1, n3, n5 ARE rack 1, so those are the
         same failure written two ways -- and both are survivable.
         Three failures only hurt when they straddle the racks, as
         (n0, n1, n3) does

    the worst case, by brute force -- how many DataNode failures
    can this layout survive with certainty?
      1 node(s) down:   0 of   6 combinations lose data
      2 node(s) down:   0 of  15 combinations lose data
      3 node(s) down:   6 of  20 combinations lose data
      4 node(s) down:  12 of  15 combinations lose data
      5 node(s) down:   6 of   6 combinations lose data
      6 node(s) down:   1 of   1 combinations lose data
         ANY TWO failures are survivable; some threes are not.
         Replication factor R tolerates R-1 arbitrary failures --
         and note that most 3-node combinations are still fine, so
         'replication 3 fails at 3 nodes' is only true of the worst
         case, which is the honest way to state it

    re-replication after a DataNode is declared dead:
      1. DataNode misses heartbeats (default: 3 sec interval)
      2. NameNode waits 10 * 3 sec + 2 * 5 min = 10 min 30 sec
      3. its blocks are now UNDER-REPLICATED (2 of 3)
      4. NameNode schedules copies from surviving replicas
      5. replication returns to 3; no client ever saw an error
         the 630-second default is deliberately LONG. A node
         that reboots in five minutes should not trigger a cluster-wide
         copy storm, so HDFS trades a longer window of reduced
         redundancy for far less needless network traffic

    the NameNode is a different kind of failure:
      component             holds                             lost on crash?
      fsimage (on disk)     the namespace at a checkpoint     no
      edit log (on disk)    changes since the checkpoint      no
      block map (in RAM)    block -> DataNode locations       YES
         the block MAP is never persisted -- it is rebuilt from
         DataNode block reports at startup, which is why a large
         NameNode takes minutes to leave safe mode. The namespace
         survives; the locations are reconstructed

    the three answers to NameNode failure, in historical order:
      mechanism                 recovers        automatic?
      Secondary NameNode        checkpoint only no -- NOT a standby
      NameNode HA (2 NNs)       full            yes, via ZooKeeper
      HDFS Federation           n/a -- scales namespacen/a
         the Secondary NameNode is the most misleadingly named
         component in Hadoop: it merges fsimage with the edit log so
         restarts stay fast, and it CANNOT take over. HA needs two
         NameNodes, a shared edit log (QJM) and ZooKeeper for failover
         -- which is exactly why experiment 16 exists

THE HONEST VERSION, BY BRUTE FORCE

Nodes down Combinations losing data
1 0 of 6
2 0 of 15
3 6 of 20
4 12 of 15
5 6 of 6

Any two failures are survivable; 14 of the 20 three-node combinations are still fine. "Replication 3 fails at 3 nodes" is the worst case, not the rule, and stating it that way is the honest answer.

THE 630-SECOND DELAY

10 × 3 s + 2 × 5 min = 630 s before a DataNode is declared dead. The delay is deliberate — a node that reboots in five minutes should not trigger a cluster-wide copy storm. The lab's cluster sets the recheck interval to 15 s, so 10 × 3 s + 2 × 15 s = 60 s, and the script waits 75 s rather than eleven minutes; the formula is the same.

THE NAMENODE'S BLOCK MAP IS NEVER PERSISTED

Component Holds Lost on crash?
fsimage (disk) the namespace at a checkpoint no
edits (disk) changes since no
block map (RAM) block → DataNode locations YES

Rebuilt from block reports at startup, which is why a large NameNode takes minutes to leave safe mode — the safemode get above caught it there. And the Secondary NameNode is a checkpointer, not a standby — the most misleadingly named component in Hadoop.

RESULT

With a DataNode killed the 300 MB file was still read in full; the node was declared dead and its blocks re-replicated to the three still alive; with the NameNode killed nothing was reachable, and after its restart and safe mode every byte was there. In the model, any two failures are survivable and 6 of the 20 three-node failures lose data.

Experiment 6 — YARN: the ResourceManager, the NodeManager and the schedulers

1. Question

Configure YARN and run sample applications, observing the roles of the ResourceManager and the NodeManager.

2. Aim

Set up two Capacity Scheduler queues, run the sample applications into each, observe the nodes and queues, and kill a running job; then compare FIFO, fair and capacity scheduling on one workload.

3. Steps

On the cluster, 06_yarn.sh:

  1. Read the configuration that matters.
  2. Set up the queues.
  3. Run the sample applications.
  4. Submit to a named queue.
  5. Observe the nodes and queues.
  6. Kill a job.

The Python check, 06_yarn_scheduling.py:

  1. Set out the workload.
  2. Schedule first in, first out.
  3. Schedule fairly.
  4. Schedule by capacity.
  5. Name who does what.

THE SCHEDULERS, ON ONE WORKLOAD

The workload: an 8-container cluster; big_etl needs all 8 for 10 s; small_q1 and small_q2 need 1 container for 2 s; medium needs 4 for 5 s.

Job FIFO Fair Capacity (75/25)
big_etl 10 14 20
small_q1 11 1 2
small_q2 12 1 3
medium 16 5 22
total turnaround 49 21 —

4. Programme

On the cluster, 06_yarn.sh:

# Experiment 6 -- configure YARN, run sample applications, observe ResourceManager and NodeManager roles
#
# Run it: bash 06_yarn.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 06_yarn_scheduling.py, which runs FIFO, Fair and Capacity on one workload
#
# Step 1: Read the configuration that matters
#   yarn.nodemanager.resource.memory-mb        8192   RAM this node offers
#   yarn.nodemanager.resource.cpu-vcores          8   cores this node offers
#   yarn.scheduler.minimum-allocation-mb       1024   container granularity
#   yarn.scheduler.maximum-allocation-mb       8192   biggest single container
#   yarn.resourcemanager.scheduler.class
#       org.apache.hadoop.yarn.server.resourcemanager.scheduler.
#       capacity.CapacityScheduler
#   yarn.nodemanager.aux-services      mapreduce_shuffle
#       ^ without this the shuffle has no server and every job hangs at 33%

# Step 2: Set up the queues
#   yarn.scheduler.capacity.root.queues                 production,adhoc
#   yarn.scheduler.capacity.root.production.capacity    75
#   yarn.scheduler.capacity.root.adhoc.capacity         25
#   yarn.scheduler.capacity.root.adhoc.maximum-capacity 50   <- ELASTICITY

yarn rmadmin -refreshQueues        # queues reload WITHOUT restarting the RM
yarn queue -status production 2>/dev/null | grep -E "Queue Name|Capacity"
yarn queue -status adhoc 2>/dev/null | grep -E "Queue Name|Capacity"

# Step 3: Run the sample applications
EX=$(ls $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar)
yarn jar $EX pi 4 1000 2>&1 | grep -E "Estimated value|Job Finished"
yarn jar $EX teragen 100000 /user/student/terasort-in 2>&1 | grep -E "completed successfully"
yarn jar $EX terasort /user/student/terasort-in /user/student/terasort-out 2>&1 | grep -E "completed successfully"
yarn jar $EX teravalidate /user/student/terasort-out /user/student/terasort-check 2>&1 | grep -E "completed successfully"
hdfs dfs -cat /user/student/terasort-check/part-r-00000
#   no "misorder" line: the output is in order. The checksum is of the rows.
# [Corrected: teragen wrote 10,000,000 rows -- a gigabyte, three on disk with
# replication 3 -- which is slow on one machine and proves nothing more than
# 100,000 rows (10 MB) do. teravalidate is added: it is what says the sort worked.]

# Step 4: Submit to a named queue
# submit into a named queue and watch where it lands
yarn jar $EX pi -Dmapreduce.job.queuename=adhoc 4 1000 2>&1 | grep -E "Estimated value"
yarn application -list -appStates FINISHED 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $5 " | " $7}' | sort
# [Corrected: this was `yarn jar ...examples-*.jar`, with three dots typed
# where the path goes; the path is set once, in EX, above.]

# Step 5: Observe the nodes and queues
# a job's client returns when the job reports success, and its last container is
# released a moment later; wait for that, so what follows is the idle cluster
until yarn node -list 2>/dev/null | grep -qE "RUNNING.*[[:space:]]0$"; do sleep 1; done
yarn node -list -all 2>/dev/null | tail -n +2  # every NodeManager, its state and containers
yarn queue -status adhoc 2>/dev/null | grep -E "Current Capacity|Maximum Capacity"

# Step 6: Kill a job
# a job to kill: start one that runs a while, and kill it by its id
yarn jar $EX pi 4 1000000000 > /dev/null 2>&1 &
sleep 15
APP=$(yarn application -list 2>/dev/null | awk '/^application_/{print $1}' | head -1)
yarn application -list 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $6}'
yarn application -kill $APP 2>&1 | grep -E "Killed application|Killing application"
wait
yarn application -status $APP 2>/dev/null | grep -E "Final-State"
# [Corrected: this killed application_1699999999999_0002, an id from someone
# else's cluster; the id is taken from the list.]
#
# yarn top                         # like top(1), for the cluster
#   -- a full-screen display that runs until you press q, so it is not run here.

#   http://localhost:8088/cluster/scheduler   the queue tree, live
#
# What to look for: submit a big job and a small one into different queues,
# and watch the adhoc job start IMMEDIATELY even though production is full.
# That guarantee is what a queue is, and it is the whole answer to
# "what does the Capacity Scheduler do".

The Python check, 06_yarn_scheduling.py:

"""Experiment 6 -- configure YARN, run sample applications, and observe the
ResourceManager and NodeManager roles.

`06_yarn.sh` carries the real commands. What runs here is the SCHEDULER, which
is the part of YARN that actually decides anything -- and the part where the
three policies give visibly different answers on the same queue.
"""

# (name, containers needed, seconds per container-slot, submitted at)
JOBS = [
    ("big_etl",  8, 10, 0),
    ("small_q1", 1,  2, 1),
    ("small_q2", 1,  2, 2),
    ("medium",   4,  5, 3),
]

CLUSTER = 8          # containers available cluster-wide


def fifo(jobs, capacity):
    """First in, first out. One job owns the cluster until it finishes."""
    now, done = 0, {}
    for name, need, secs, submitted in sorted(jobs, key=lambda j: j[3]):
        start = max(now, submitted)
        waves = -(-need // capacity)          # ceil
        now = start + waves * secs
        done[name] = (start, now, now - submitted)
    return done


def fair(jobs, capacity):
    """Fair scheduler: every RUNNING job gets an equal share of containers.

    Simulated one second at a time, which is crude but exactly right for
    showing the property that matters -- a one-container job does not wait
    behind an eight-container job.
    """
    remaining = {j[0]: j[1] * j[2] for j in jobs}   # container-seconds of work
    submitted = {j[0]: j[3] for j in jobs}
    done, t = {}, 0
    while any(v > 0 for v in remaining.values()):
        active = [n for n, v in remaining.items() if v > 0 and submitted[n] <= t]
        if not active:
            t += 1
            continue
        share = capacity / len(active)
        for n in active:
            remaining[n] -= share
            if remaining[n] <= 0 and n not in done:
                done[n] = (submitted[n], t + 1, t + 1 - submitted[n])
        t += 1
    return done


def capacity_sched(jobs, capacity, queues):
    """Capacity scheduler: queues get guaranteed percentages of the cluster.

    A job cannot exceed its queue's share even when the cluster is idle,
    unless elasticity is enabled -- which is the whole difference between
    'capacity' and 'fair'.
    """
    done = {}
    for qname, pct, members in queues:
        slots = max(1, int(capacity * pct / 100))
        qjobs = [j for j in jobs if j[0] in members]
        sub = fifo(qjobs, slots)
        for k, v in sub.items():
            done[k] = v
    return done


def main():
    print("  Experiment 6 -- YARN scheduling")

    # Step 1: Set out the workload
    print("\n    the workload, on an 8-container cluster:")
    print(f"      {'job':<10}{'containers':>11}{'sec/wave':>10}{'submitted':>11}")
    for name, need, secs, sub in JOBS:
        print(f"      {name:<10}{need:>11}{secs:>10}{sub:>11}")

    f = fifo(JOBS, CLUSTER)
    # Step 2: Schedule first in, first out
    print("\n    FIFO scheduler:")
    print(f"      {'job':<10}{'start':>7}{'finish':>8}{'turnaround':>12}")
    for name, _, _, _ in JOBS:
        s, e, t = f[name]
        print(f"      {name:<10}{s:>7}{e:>8}{t:>12}")
    fifo_small = f["small_q1"][2]
    print(f"""         small_q1 needs ONE container for TWO seconds and waits
         {fifo_small} seconds, because big_etl took the whole cluster first.
         That is head-of-line blocking, and it is why nobody runs
         FIFO on a shared cluster""")

    fr = fair(JOBS, CLUSTER)
    # Step 3: Schedule fairly
    print("\n    Fair scheduler:")
    print(f"      {'job':<10}{'start':>7}{'finish':>8}{'turnaround':>12}")
    for name, _, _, _ in JOBS:
        s, e, t = fr[name]
        print(f"      {name:<10}{s:>7}{e:>8}{t:>12}")
    fair_small = fr["small_q1"][2]
    assert fair_small < fifo_small
    print(f"""         small_q1 now finishes in {fair_small}s instead of {fifo_small}s.
         Fair sharing did not make the cluster faster -- big_etl
         finished LATER ({f['big_etl'][1]} -> {fr['big_etl'][1]}) -- it moved latency from
         the small job to the big one, which is almost always the
         trade you want on an interactive cluster""")

    total_fifo = sum(v[2] for v in f.values())
    total_fair = sum(v[2] for v in fr.values())
    work = sum(need * secs for _, need, secs, _ in JOBS)
    print(f"\n    total container-seconds of WORK: {work} either way")
    print(f"    total turnaround:  FIFO {total_fifo}   Fair {total_fair}")
    assert total_fair < total_fifo
    print(f"""         the work is identical -- {work} container-seconds, which on
         8 containers cannot finish before second {-(-work // CLUSTER)}. What changed is
         WAITING: FIFO made three jobs queue behind one, so total
         turnaround fell from {total_fifo} to {total_fair} without the cluster doing
         anything faster.
         Scheduling decides WHO waits. It cannot create throughput,
         but idle-while-queued is real waste and fair sharing removes
         it""")

    cap = capacity_sched(JOBS, CLUSTER, [
        ("production", 75, {"big_etl", "medium"}),
        ("adhoc",      25, {"small_q1", "small_q2"}),
    ])
    # Step 4: Schedule by capacity
    print("\n    Capacity scheduler -- production 75%, adhoc 25%:")
    print(f"      {'job':<10}{'queue':<12}{'start':>7}{'finish':>8}{'turnaround':>12}")
    for name, q in (("big_etl", "production"), ("medium", "production"),
                    ("small_q1", "adhoc"), ("small_q2", "adhoc")):
        s, e, t = cap[name]
        print(f"      {name:<10}{q:<12}{s:>7}{e:>8}{t:>12}")
    print("""         the adhoc queue holds 2 containers whatever else is
         running, so a short query has a GUARANTEE rather than a
         hope. The cost: those 2 containers sit idle when adhoc is
         empty, unless queue elasticity is turned on""")

    # Step 5: Name who does what
    print("\n    who does what in YARN:")
    print(f"      {'component':<22}{'one per':<14}{'responsibility'}")
    for c, per, resp in (
            ("ResourceManager", "cluster", "global scheduling; hands out containers"),
            ("NodeManager", "node", "launches and monitors containers, reports health"),
            ("ApplicationMaster", "JOB", "negotiates containers, retries failed tasks"),
            ("Container", "task", "a bounded slice of CPU and RAM on one node")):
        print(f"      {c:<22}{per:<14}{resp}")
    print("""         ONE ApplicationMaster PER JOB is the change that defined
         YARN. In Hadoop 1 the JobTracker did both scheduling and job
         management for every job, so it was the bottleneck AND the
         single point of failure. Splitting them is why YARN can run
         Spark, Tez and Flink and not only MapReduce""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 06_yarn.sh:

OUTPUT

$ yarn rmadmin -refreshQueues
2026-10-04 23:12:43,788 INFO client.DefaultNoHARMFailoverProxyProvider: Connecting to ResourceManager at /0.0.0.0:8033
$ yarn queue -status production 2>/dev/null | grep -E "Queue Name|Capacity"
Queue Name : production
    Capacity : 75.00%
    Current Capacity : .00%
    Maximum Capacity : 100.00%
$ yarn queue -status adhoc 2>/dev/null | grep -E "Queue Name|Capacity"
Queue Name : adhoc
    Capacity : 25.00%
    Current Capacity : .00%
    Maximum Capacity : 50.00%
$ EX=$(ls $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar)
$ yarn jar $EX pi 4 1000 2>&1 | grep -E "Estimated value|Job Finished"
Job Finished in 24.944 seconds
Estimated value of Pi is 3.14000000000000000000
$ yarn jar $EX teragen 100000 /user/student/terasort-in 2>&1 | grep -E "completed successfully"
2026-10-04 23:13:29,432 INFO mapreduce.Job: Job job_1791155559313_0002 completed successfully
$ yarn jar $EX terasort /user/student/terasort-in /user/student/terasort-out 2>&1 | grep -E "completed successfully"
2026-10-04 23:13:51,600 INFO mapreduce.Job: Job job_1791155559313_0003 completed successfully
$ yarn jar $EX teravalidate /user/student/terasort-out /user/student/terasort-check 2>&1 | grep -E "completed successfully"
2026-10-04 23:14:11,151 INFO mapreduce.Job: Job job_1791155559313_0004 completed successfully
$ hdfs dfs -cat /user/student/terasort-check/part-r-00000
checksum    c327a1c42c28
$ yarn jar $EX pi -Dmapreduce.job.queuename=adhoc 4 1000 2>&1 | grep -E "Estimated value"
Estimated value of Pi is 3.14000000000000000000
$ yarn application -list -appStates FINISHED 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $5 " | " $7}' | sort
             TeraGen | production |          SUCCEEDED
            TeraSort | production |          SUCCEEDED
        TeraValidate | production |          SUCCEEDED
     QuasiMonteCarlo |      adhoc |          SUCCEEDED
     QuasiMonteCarlo | production |          SUCCEEDED
$ until yarn node -list 2>/dev/null | grep -qE "RUNNING.*[[:space:]]0$"; do sleep 1; done
$ yarn node -list -all 2>/dev/null | tail -n +2
         Node-Id         Node-State Node-Http-Address   Number-of-Running-Containers
 localhost:36875            RUNNING    localhost:8042                              0
$ yarn queue -status adhoc 2>/dev/null | grep -E "Current Capacity|Maximum Capacity"
    Current Capacity : .00%
    Maximum Capacity : 50.00%
$ yarn jar $EX pi 4 1000000000 > /dev/null 2>&1 &
$ sleep 15
$ APP=$(yarn application -list 2>/dev/null | awk '/^application_/{print $1}' | head -1)
$ yarn application -list 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $6}'
     QuasiMonteCarlo |            RUNNING
$ yarn application -kill $APP 2>&1 | grep -E "Killed application|Killing application"
Killing application application_1791155559313_0006
2026-10-04 23:15:07,634 INFO impl.YarnClientImpl: Killed application application_1791155559313_0006
$ wait
$ yarn application -status $APP 2>/dev/null | grep -E "Final-State"
    Final-State : KILLED

The Python check, 06_yarn_scheduling.py:

OUTPUT

  Experiment 6 -- YARN scheduling

    the workload, on an 8-container cluster:
      job        containers  sec/wave  submitted
      big_etl             8        10          0
      small_q1            1         2          1
      small_q2            1         2          2
      medium              4         5          3

    FIFO scheduler:
      job         start  finish  turnaround
      big_etl         0      10          10
      small_q1       10      12          11
      small_q2       12      14          12
      medium         14      19          16
         small_q1 needs ONE container for TWO seconds and waits
         11 seconds, because big_etl took the whole cluster first.
         That is head-of-line blocking, and it is why nobody runs
         FIFO on a shared cluster

    Fair scheduler:
      job         start  finish  turnaround
      big_etl         0      14          14
      small_q1        1       2           1
      small_q2        2       3           1
      medium          3       8           5
         small_q1 now finishes in 1s instead of 11s.
         Fair sharing did not make the cluster faster -- big_etl
         finished LATER (10 -> 14) -- it moved latency from
         the small job to the big one, which is almost always the
         trade you want on an interactive cluster

    total container-seconds of WORK: 104 either way
    total turnaround:  FIFO 49   Fair 21
         the work is identical -- 104 container-seconds, which on
         8 containers cannot finish before second 13. What changed is
         WAITING: FIFO made three jobs queue behind one, so total
         turnaround fell from 49 to 21 without the cluster doing
         anything faster.
         Scheduling decides WHO waits. It cannot create throughput,
         but idle-while-queued is real waste and fair sharing removes
         it

    Capacity scheduler -- production 75%, adhoc 25%:
      job       queue         start  finish  turnaround
      big_etl   production        0      20          20
      medium    production       20      25          22
      small_q1  adhoc             1       3           2
      small_q2  adhoc             3       5           3
         the adhoc queue holds 2 containers whatever else is
         running, so a short query has a GUARANTEE rather than a
         hope. The cost: those 2 containers sit idle when adhoc is
         empty, unless queue elasticity is turned on

    who does what in YARN:
      component             one per       responsibility
      ResourceManager       cluster       global scheduling; hands out containers
      NodeManager           node          launches and monitors containers, reports health
      ApplicationMaster     JOB           negotiates containers, retries failed tasks
      Container             task          a bounded slice of CPU and RAM on one node
         ONE ApplicationMaster PER JOB is the change that defined
         YARN. In Hadoop 1 the JobTracker did both scheduling and job
         management for every job, so it was the bottleneck AND the
         single point of failure. Splitting them is why YARN can run
         Spark, Tez and Flink and not only MapReduce

WHAT THE NUMBERS SAY, STATED CAREFULLY

THE CAPACITY SCHEDULER'S GUARANTEE

The adhoc queue holds 25% of the cluster whatever else is running, so a short query has a guarantee rather than a hope — the queues the script set up report 75% and 25%, with adhoc allowed to grow to 50%. The cost: that share sits idle when adhoc is empty, unless maximum-capacity is raised to allow elasticity — which is exactly the difference between "capacity" and "fair".

One ApplicationMaster per job is the change that defined YARN. Hadoop 1's JobTracker did both scheduling and per-job management, so it was the bottleneck and the single point of failure, and it could run only MapReduce.

RESULT

Pi, TeraGen, TeraSort and TeraValidate ran in the production queue and Pi again in adhoc, each SUCCEEDED; a long Pi job was killed by its id and finished KILLED. In the model, fair sharing halves total turnaround, 49 s to 21 s, on the same 104 container-seconds of work.

Experiment 7 — Word count in MapReduce

1. Question

Write a simple MapReduce program for word count.

2. Aim

Count the words in six documents with a Java MapReduce job, with a combiner; then make the shuffle visible with a MapReduce engine in Python, and see when a combiner is safe.

3. Steps

On the cluster, WordCount.java:

  1. Map each word to 1.
  2. Reduce by summing.
  3. Configure the job, with a combiner.

The Python check, 07_wordcount.py:

  1. Map, shuffle and reduce the documents.
  2. Read the top words.
  3. Add a combiner.
  4. See when a combiner is not safe.
  5. Partition the keys.

THE SHUFFLE, MADE VISIBLE

mapreduce.py, which 07_wordcount.py imports, is a MapReduce engine in forty lines, written out in full. The point is that it makes the shuffle visible, and the shuffle is the part students never see and the part that costs the money.

Phase Records
map output 48
shuffled 48
reduce output 26

Top words: the 5, big 4, data 4, dog 4, quick 3, fox 3. The counts sum back to 48 — reduce is a regrouping, and if your totals do not reconcile, your reducer is not associative.

4. Programme

On the cluster, WordCount.java:

// Experiment 7 -- word count in MapReduce
//
// Run it: the build-and-run lines below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 07_wordcount.py, which runs the same map and reduce
// functions through an explicit engine and asserts every count
//
// Build and run:
//   javac -classpath $(hadoop classpath) -d classes WordCount.java
//   jar -cvf WordCount.jar -C classes/ .
//   hadoop jar WordCount.jar WordCount /user/student/docs /user/student/out
//

import java.io.IOException;
import java.util.StringTokenizer;

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

public class WordCount {

    // Step 1: Map each word to 1
    public static class TokenizerMapper
            extends Mapper<LongWritable, Text, Text, IntWritable> {

        // Reused across every call. Allocating a new Text per word would
        // create one object per word in the corpus, and the GC pause is the
        // job. This is the single most important idiom in MapReduce Java.
        private final static IntWritable ONE = new IntWritable(1);
        private final Text word = new Text();

        @Override
        public void map(LongWritable key, Text value, Context context)
                throws IOException, InterruptedException {
            // key is the BYTE OFFSET of the line, not a line number.
            StringTokenizer itr = new StringTokenizer(value.toString());
            while (itr.hasMoreTokens()) {
                word.set(itr.nextToken().toLowerCase());
                context.write(word, ONE);
            }
        }
    }

    // Step 2: Reduce by summing
    public static class IntSumReducer
            extends Reducer<Text, IntWritable, Text, IntWritable> {

        private final IntWritable result = new IntWritable();

        @Override
        public void reduce(Text key, Iterable<IntWritable> values, Context context)
                throws IOException, InterruptedException {
            int sum = 0;
            for (IntWritable val : values) {
                sum += val.get();
            }
            // The Iterable is streamed from disk and can be walked ONCE.
            // Calling values.iterator() a second time yields nothing -- the
            // classic bug when someone tries to compute a mean and a count.
            result.set(sum);
            context.write(key, result);
        }
    }

    // Step 3: Configure the job, with a combiner
    public static void main(String[] args) throws Exception {
        Configuration conf = new Configuration();
        Job job = Job.getInstance(conf, "word count");
        job.setJarByClass(WordCount.class);

        job.setMapperClass(TokenizerMapper.class);
        job.setCombinerClass(IntSumReducer.class);   // safe: sum is associative
        job.setReducerClass(IntSumReducer.class);

        job.setOutputKeyClass(Text.class);
        job.setOutputValueClass(IntWritable.class);
        job.setNumReduceTasks(1);

        FileInputFormat.addInputPath(job, new Path(args[0]));
        FileOutputFormat.setOutputPath(job, new Path(args[1]));
        // The output directory MUST NOT EXIST. Hadoop refuses to overwrite,
        // which prevents a re-run from silently destroying yesterday's result.

        System.exit(job.waitForCompletion(true) ? 0 : 1);
    }
}

// Expected output on the six sample documents (verified in 07_wordcount.py):
//   the 5,  big 4,  data 4,  dog 4,  fox 3,  quick 3,  ... 26 terms, 48 words
//
// The combiner is the same class as the reducer here ONLY because sum is
// associative and commutative. Setting a combiner for a mean silently
// produces wrong answers -- mean of means is not the mean -- and Hadoop will
// not warn you.

The Python check, 07_wordcount.py:

"""Experiment 7 -- a simple MapReduce program for word count.

This is the "hello world" of MapReduce, and it is worth more than it looks:
the shuffle between map and reduce is the only part of the model that costs
real money on a cluster, and word count is the smallest program that makes it
visible.

Runs through the engine in mapreduce.py, and then through REAL PYSPARK if the
Spark virtual environment is present (see tools/setup_spark.sh).
"""
from mapreduce import run
import fixtures as f

INPUT = sorted(f.DOCS.items())


def mapper(name, line):
    """(filename, line) -> (word, 1) for every word."""
    for word in line.split():
        yield word, 1


def reducer(word, counts):
    """(word, [1, 1, ...]) -> (word, total)."""
    yield word, sum(counts)


def main():
    print("  Experiment 7 -- word count in MapReduce")

    print(f"\n    input: {len(INPUT)} documents, "
          f"{sum(len(t.split()) for _, t in INPUT)} words")

    # Step 1: Map, shuffle and reduce the documents
    trace = {}
    result = run(INPUT, mapper, reducer, trace=trace)
    counts = dict(result)

    print(f"\n    {'phase':<26}{'records':>9}")
    print(f"    {'map output':<26}{trace['map_output']:>9}")
    print(f"    {'shuffled across network':<26}{trace['shuffled']:>9}")
    print(f"    {'reduce output':<26}{len(result):>9}")
    assert trace["map_output"] == 48
    assert len(result) == 26

    # Step 2: Read the top words
    print("\n    the top words:")
    for w, c in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))[:6]:
        print(f"      {w:<10}{c:>3}")
    assert counts["the"] == 5 and counts["dog"] == 4
    assert counts["big"] == 4 and counts["data"] == 4
    assert sum(counts.values()) == 48
    print("""         the counts sum back to 48, the map output. Nothing was
         created or lost -- reduce is a REGROUPING, and if your
         totals do not reconcile, your reducer is not associative""")

    # Step 3: Add a combiner
    ctrace = {}
    combined = run(INPUT, mapper, reducer, combiner=reducer, trace=ctrace)
    assert combined == result, "a combiner must not change the answer"
    saved = trace["shuffled"] - ctrace["shuffled"]
    pct = 100 * saved / trace["shuffled"]
    print(f"\n    with a combiner (the reducer, run map-side, PER TASK):")
    print(f"      shuffled {trace['shuffled']} -> {ctrace['shuffled']} "
          f"({saved} fewer records, {pct:.2f}%)")
    print(f"""         same answer, {pct:.1f}% less network. And note how SMALL that
         saving is: these documents are 5 to 11 words, so there is
         almost nothing to merge within one split. On a 128 MB split
         of real text the same combiner cuts the shuffle by orders of
         magnitude. The combiner's value scales with SPLIT SIZE, which
         is the point this tiny dataset makes by failing to impress""")

    # Step 4: See when a combiner is not safe
    print("\n    when a combiner is NOT safe:")
    print(f"      {'reducer computes':<22}{'combiner-safe?':<16}why")
    for what, safe, why in (
            ("sum", "yes", "associative and commutative"),
            ("max", "yes", "max of maxes is the max"),
            ("count", "yes", "if the combiner emits partial counts"),
            ("MEAN", "NO", "mean of means is not the mean"),
            ("median", "NO", "needs every value at once")):
        print(f"      {what:<22}{safe:<16}{why}")
    # prove the mean case rather than asserting it
    groups = [[1, 1, 1, 10], [10]]
    naive = sum(sum(g) / len(g) for g in groups) / len(groups)
    true = sum(sum(g) for g in groups) / sum(len(g) for g in groups)
    print(f"\n      mean of means = {naive:.4f}, true mean = {true:.4f}")
    assert abs(naive - true) > 1
    print(f"""         {naive:.4f} against {true:.4f} on five numbers. To average safely,
         emit (sum, count) pairs from the combiner and divide only in
         the reducer -- and that is the same average-of-averages trap
         Course 11 met in DAX, in a different costume""")

    # Step 5: Partition the keys
    print("\n    3 reduce tasks instead of 1:")
    ptrace = {}
    three = run(INPUT, mapper, reducer, reducers=3, trace=ptrace)
    assert three == result, "the number of reducers must not change the answer"
    sizes = ptrace["partition_sizes"]
    print(f"      partition sizes: {sizes}   (total {sum(sizes)})")
    print(f"      largest / smallest = {max(sizes) / min(sizes):.2f}")
    print("""         hash partitioning is only as balanced as the KEY
         DISTRIBUTION. Natural language is Zipfian, so a real corpus
         skews far worse than this -- one reducer gets 'the' and
         finishes last, and the job's wall clock is that reducer.
         Skew, not volume, is what usually kills a MapReduce job""")

    return counts


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, WordCount.java:

OUTPUT

$ hdfs dfs -mkdir -p /user/student/docs
$ hdfs dfs -put docs/*.txt /user/student/docs/
$ hdfs dfs -ls /user/student/docs | awk 'NR>1{print $NF}'
/user/student/docs/doc1.txt
/user/student/docs/doc2.txt
/user/student/docs/doc3.txt
/user/student/docs/doc4.txt
/user/student/docs/doc5.txt
/user/student/docs/doc6.txt
$ mkdir -p classes
$ javac -classpath $(hadoop classpath) -d classes WordCount.java
$ jar -cvf WordCount.jar -C classes/ .
added manifest
adding: WordCount$TokenizerMapper.class(in = 1895) (out= 799)(deflated 57%)
adding: WordCount.class(in = 1530) (out= 832)(deflated 45%)
adding: WordCount$IntSumReducer.class(in = 1739) (out= 740)(deflated 57%)
$ hadoop jar WordCount.jar WordCount /user/student/docs /user/student/out 2>&1 | grep -E "completed successfully|Map input records|Map output records|Combine input records|Combine output records|Reduce input groups|Reduce output records"
2026-10-04 23:16:12,108 INFO mapreduce.Job: Job job_1791155734078_0001 completed successfully
        Map input records=6
        Map output records=48
        Combine input records=48
        Combine output records=39
        Reduce input groups=26
        Reduce output records=26
$ hdfs dfs -cat /user/student/out/part-r-00000
a   2
all 1
and 2
big 4
brown   2
data    4
day 1
dog 4
for 1
fox 3
hadoop  1
is  2
jumps   1
lazy    2
machine 1
one 1
outpaces    1
over    1
processes   1
quick   3
sleeps  1
spark   1
stores  1
that    1
the 5
too 1

The Python check, 07_wordcount.py:

OUTPUT

  Experiment 7 -- word count in MapReduce

    input: 6 documents, 48 words

    phase                       records
    map output                       48
    shuffled across network          48
    reduce output                    26

    the top words:
      the         5
      big         4
      data        4
      dog         4
      fox         3
      quick       3
         the counts sum back to 48, the map output. Nothing was
         created or lost -- reduce is a REGROUPING, and if your
         totals do not reconcile, your reducer is not associative

    with a combiner (the reducer, run map-side, PER TASK):
      shuffled 48 -> 39 (9 fewer records, 18.75%)
         same answer, 18.8% less network. And note how SMALL that
         saving is: these documents are 5 to 11 words, so there is
         almost nothing to merge within one split. On a 128 MB split
         of real text the same combiner cuts the shuffle by orders of
         magnitude. The combiner's value scales with SPLIT SIZE, which
         is the point this tiny dataset makes by failing to impress

    when a combiner is NOT safe:
      reducer computes      combiner-safe?  why
      sum                   yes             associative and commutative
      max                   yes             max of maxes is the max
      count                 yes             if the combiner emits partial counts
      MEAN                  NO              mean of means is not the mean
      median                NO              needs every value at once

      mean of means = 6.6250, true mean = 4.6000
         6.6250 against 4.6000 on five numbers. To average safely,
         emit (sum, count) pairs from the combiner and divide only in
         the reducer -- and that is the same average-of-averages trap
         Course 11 met in DAX, in a different costume

    3 reduce tasks instead of 1:
      partition sizes: [20, 17, 11]   (total 48)
      largest / smallest = 1.82
         hash partitioning is only as balanced as the KEY
         DISTRIBUTION. Natural language is Zipfian, so a real corpus
         skews far worse than this -- one reducer gets 'the' and
         finishes last, and the job's wall clock is that reducer.
         Skew, not volume, is what usually kills a MapReduce job

THE COMBINER, AND WHY ITS SAVING IS SMALL HERE

Shuffled
no combiner 48
with combiner 39
saving 9 (18.75%)

WHY IT MATTERS

Note how small that is, and why it is honest. The combiner runs per map task, and these documents are 5 to 11 words — there is almost nothing to merge within one split. On a 128 MB split the same combiner cuts the shuffle by orders of magnitude. The cluster agrees: its counters show the same 48 map output records and 39 combine output records.

The combiner's value scales with split size, which is the point this tiny dataset makes precisely by failing to impress.

THE COMBINER THAT IS WRONG

Reducer computes Safe?
sum, max, count yes
mean NO
median NO

Demonstrated: [1,1,1,10] and [10] give mean of means 6.6250 against true mean 4.6000. Emit (sum, count) and divide only in the reducer — the same average-of-averages trap Business Intelligence Tools met in DAX.

Partitioning. 3 reducers give partitions of [20, 17, 11], ratio 1.82. Hash partitioning is only as balanced as the key distribution, and skew, not volume, is what usually kills a MapReduce job.

The three Java details. The map key is a byte offset, not a line number. Reuse the Writable objects — a new Text() per word makes GC the job. The reduce Iterable can be walked once — it streams from disk.

RESULT

48 words in, 26 distinct words out, on the cluster and in the engine alike; the combiner cut the 48 shuffled records to 39 — 18.75%, small because each split is one short document. A combiner for a mean is wrong: 6.6250 against the true 4.6000.

Experiment 8 — An inverted index in MapReduce

1. Question

Develop a MapReduce job for inverted index creation.

2. Aim

Build, with a Java MapReduce job, an index from each word to the documents it occurs in and how often; then answer queries from the index alone, and measure its size and skew.

3. Steps

On the cluster, InvertedIndex.java:

  1. Map each word to its document.
  2. Reduce to a posting list.
  3. Configure the job.

The Python check, 08_inverted_index.py:

  1. Build the index.
  2. Read a slice of it.
  3. Answer queries from it.
  4. Weigh its size.
  5. See why it is a MapReduce job.
  6. Measure the skew.

THE INDEX

Term Postings
dog doc1:1, doc2:1, doc3:1, doc6:1
quick doc1:1, doc3:2
big doc4:2, doc5:2

quick appears twice in doc3, and the posting records it. Frequency is the difference between "does this word occur" and "how relevant is this document" — boolean retrieval against ranked retrieval, in one number.

4. Programme

On the cluster, InvertedIndex.java:

// Experiment 8 -- an inverted index in MapReduce
//
// Run it: the build-and-run lines below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 08_inverted_index.py, which builds the same index and
// answers boolean queries against it
//
// Build and run:
//   javac -classpath $(hadoop classpath) -d classes InvertedIndex.java
//   jar -cvf InvertedIndex.jar -C classes/ .
//   hadoop jar InvertedIndex.jar InvertedIndex /user/student/docs /user/student/out
//

import java.io.IOException;
import java.util.HashMap;
import java.util.Map;
import java.util.StringTokenizer;
import java.util.TreeMap;

import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.input.FileSplit;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;

public class InvertedIndex {

    // Step 1: Map each word to its document
    public static class IndexMapper
            extends Mapper<LongWritable, Text, Text, Text> {

        private final Text word = new Text();
        private final Text docId = new Text();
        private String fileName;

        @Override
        protected void setup(Context context) {
            // THE FILENAME IS NOT IN THE KEY OR THE VALUE. It comes from the
            // InputSplit, and this is the only way to get at it. Every
            // inverted-index question turns on knowing that.
            FileSplit split = (FileSplit) context.getInputSplit();
            fileName = split.getPath().getName();
        }

        @Override
        public void map(LongWritable key, Text value, Context context)
                throws IOException, InterruptedException {
            StringTokenizer itr = new StringTokenizer(value.toString());
            while (itr.hasMoreTokens()) {
                word.set(itr.nextToken().toLowerCase());
                docId.set(fileName);
                context.write(word, docId);
            }
        }
    }

    // Step 2: Reduce to a posting list
    public static class IndexReducer
            extends Reducer<Text, Text, Text, Text> {

        private final Text postings = new Text();

        @Override
        public void reduce(Text key, Iterable<Text> values, Context context)
                throws IOException, InterruptedException {
            // Count occurrences per document, in ONE pass over the Iterable.
            Map<String, Integer> freq = new TreeMap<>();
            for (Text doc : values) {
                String d = doc.toString();
                freq.merge(d, 1, Integer::sum);
            }
            StringBuilder sb = new StringBuilder();
            for (Map.Entry<String, Integer> e : freq.entrySet()) {
                if (sb.length() > 0) sb.append(", ");
                sb.append(e.getKey()).append(':').append(e.getValue());
            }
            postings.set(sb.toString());
            context.write(key, postings);
            // The posting list is held in memory here. For a stop word on a
            // real corpus that map does not fit, which is why production
            // indexers emit (term, doc) pairs SORTED and stream the merge --
            // a secondary sort, not a HashMap.
        }
    }

    // Step 3: Configure the job
    public static void main(String[] args) throws Exception {
        Configuration conf = new Configuration();
        Job job = Job.getInstance(conf, "inverted index");
        job.setJarByClass(InvertedIndex.class);

        job.setMapperClass(IndexMapper.class);
        job.setReducerClass(IndexReducer.class);
        // NO COMBINER HERE. The reducer's output type (Text postings) differs
        // from its input type (Text docId), and a combiner must have the same
        // input and output types as the reducer. Setting one would not
        // compile -- and if the types happened to match, it would corrupt
        // the counts.

        job.setOutputKeyClass(Text.class);
        job.setOutputValueClass(Text.class);

        FileInputFormat.addInputPath(job, new Path(args[0]));
        FileOutputFormat.setOutputPath(job, new Path(args[1]));

        System.exit(job.waitForCompletion(true) ? 0 : 1);
    }
}

// Expected output on the six sample documents (verified in 08_inverted_index.py):
//   dog      doc1.txt:1, doc2.txt:1, doc3.txt:1, doc6.txt:1
//   quick    doc1.txt:1, doc3.txt:2
//   big      doc4.txt:2, doc5.txt:2
//   ... 26 terms, 39 postings

The Python check, 08_inverted_index.py:

"""Experiment 8 -- a MapReduce job for inverted index creation.

The inverted index is what a search engine is. Word count changes the VALUE
between map and reduce; the inverted index changes the KEY SPACE -- the map
output key is a word and the value is a document, which is exactly the
transposition a search index needs.
"""
from mapreduce import run
import fixtures as f

INPUT = sorted(f.DOCS.items())


def mapper(doc, text):
    """(doc, text) -> (word, doc) -- emit the document as the VALUE."""
    for pos, word in enumerate(text.split()):
        yield word, (doc, pos)


def reducer(word, postings):
    """(word, [(doc, pos), ...]) -> (word, posting list).

    The posting list is deduplicated by document and carries a frequency,
    which is what turns a boolean index into a ranked one.
    """
    per_doc = {}
    for doc, pos in postings:
        per_doc.setdefault(doc, []).append(pos)
    yield word, {d: len(p) for d, p in sorted(per_doc.items())}


def boolean_and(index, *words):
    """Intersect posting lists -- an AND query."""
    sets = [set(index.get(w, {})) for w in words]
    return sorted(set.intersection(*sets)) if sets else []


def boolean_or(index, *words):
    sets = [set(index.get(w, {})) for w in words]
    return sorted(set.union(*sets)) if sets else []


def main():
    print("  Experiment 8 -- inverted index in MapReduce")

    # Step 1: Build the index
    trace = {}
    index = dict(run(INPUT, mapper, reducer, trace=trace))

    print(f"\n    {len(INPUT)} documents in, {len(index)} index terms out")
    print(f"    map emitted {trace['map_output']} postings")
    assert trace["map_output"] == 48 and len(index) == 26

    # Step 2: Read a slice of it
    print("\n    a slice of the index (term -> {doc: frequency}):")
    for w in ("dog", "quick", "big", "data", "machine"):
        entry = ", ".join(f"{d.replace('.txt', '')}:{c}"
                          for d, c in index[w].items())
        print(f"      {w:<10}{entry}")
    assert index["dog"] == {"doc1.txt": 1, "doc2.txt": 1,
                            "doc3.txt": 1, "doc6.txt": 1}
    assert index["quick"] == {"doc1.txt": 1, "doc3.txt": 2}
    print("""         'quick' appears TWICE in doc3, and the posting records
         that. Frequency is the difference between 'does this word
         occur' and 'how relevant is this document' -- boolean
         retrieval against ranked retrieval, in one number""")

    # Step 3: Answer queries from it
    print("\n    queries answered from the index alone:")
    for q in (("quick", "fox"), ("big", "data"), ("dog", "machine")):
        hits = boolean_and(index, *q)
        pretty = [h.replace(".txt", "") for h in hits]
        print(f"      {' AND '.join(q):<24}-> {pretty if pretty else 'no match'}")
    assert boolean_and(index, "quick", "fox") == ["doc1.txt", "doc3.txt"]
    assert boolean_and(index, "big", "data") == ["doc4.txt", "doc5.txt"]
    assert boolean_and(index, "dog", "machine") == []
    hits_or = boolean_or(index, "dog", "machine")
    print(f"      {'dog OR machine':<24}-> "
          f"{[h.replace('.txt', '') for h in hits_or]}")
    assert len(hits_or) == 5
    print("""         NOT ONE DOCUMENT WAS READ to answer these. That is the
         entire point of an inverted index: query cost depends on the
         number of MATCHES, not on the size of the corpus. Scanning
         6 documents is cheap; scanning 6 billion is not""")

    # Step 4: Weigh its size
    corpus_chars = sum(len(t) for _, t in INPUT)
    postings = sum(len(v) for v in index.values())
    print(f"\n    the index is not free:")
    print(f"      corpus       {corpus_chars:>5} characters")
    print(f"      index terms  {len(index):>5}")
    print(f"      postings     {postings:>5}")
    print(f"      ratio        {postings / len(INPUT):>5.1f} postings per document")
    print("""         a full-text index typically runs 20-40% of the corpus
         size, and that is BEFORE positions. Search is a space-for-time
         trade, and 'the index is bigger than I expected' is the normal
         outcome, not a mistake""")

    # Step 5: See why it is a MapReduce job
    print("\n    why MapReduce suits this:")
    print("      map    is per-document and EMBARRASSINGLY PARALLEL")
    print("      reduce is per-term, and every posting for a term")
    print("             arrives at the same reducer by construction")
    print("""         the shuffle does the hard part -- gathering every
         mention of a word from every machine in the cluster -- and
         you never wrote a line of network code. That is the whole
         value proposition of the model""")

    # Step 6: Measure the skew
    print("\n    the skew, measured:")
    sizes = sorted(((w, sum(v.values())) for w, v in index.items()),
                   key=lambda kv: -kv[1])
    print(f"      largest posting list : {sizes[0][0]!r} with {sizes[0][1]}")
    print(f"      singleton terms      : "
          f"{sum(1 for _, n in sizes if n == 1)} of {len(sizes)}")
    assert sizes[0][0] == "the" and sizes[0][1] == 5
    print("""         'the' is the biggest list here and would be the biggest
         on any English corpus. Real engines drop stop words or split
         hot terms across reducers, because one reducer holding 'the'
         is the job's critical path""")

    return index


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, InvertedIndex.java:

OUTPUT

$ hdfs dfs -mkdir -p /user/student/docs
$ hdfs dfs -put docs/*.txt /user/student/docs/
$ hdfs dfs -ls /user/student/docs | awk 'NR>1{print $NF}'
/user/student/docs/doc1.txt
/user/student/docs/doc2.txt
/user/student/docs/doc3.txt
/user/student/docs/doc4.txt
/user/student/docs/doc5.txt
/user/student/docs/doc6.txt
$ mkdir -p classes
$ javac -classpath $(hadoop classpath) -d classes InvertedIndex.java
$ jar -cvf InvertedIndex.jar -C classes/ .
added manifest
adding: InvertedIndex$IndexReducer.class(in = 3129) (out= 1331)(deflated 57%)
adding: InvertedIndex$IndexMapper.class(in = 2342) (out= 935)(deflated 60%)
adding: InvertedIndex.class(in = 1424) (out= 772)(deflated 45%)
$ hadoop jar InvertedIndex.jar InvertedIndex /user/student/docs /user/student/out 2>&1 | grep -E "completed successfully|Map input records|Map output records|Combine input records|Combine output records|Reduce input groups|Reduce output records"
2026-10-04 23:17:18,259 INFO mapreduce.Job: Job job_1791155800966_0001 completed successfully
        Map input records=6
        Map output records=48
        Combine input records=0
        Combine output records=0
        Reduce input groups=26
        Reduce output records=26
$ hdfs dfs -cat /user/student/out/part-r-00000
a   doc3.txt:2
all doc2.txt:1
and doc5.txt:1, doc6.txt:1
big doc4.txt:2, doc5.txt:2
brown   doc1.txt:1, doc3.txt:1
data    doc4.txt:2, doc5.txt:2
day doc2.txt:1
dog doc1.txt:1, doc2.txt:1, doc3.txt:1, doc6.txt:1
for doc4.txt:1
fox doc1.txt:1, doc3.txt:1, doc6.txt:1
hadoop  doc5.txt:1
is  doc4.txt:2
jumps   doc1.txt:1
lazy    doc1.txt:1, doc2.txt:1
machine doc4.txt:1
one doc4.txt:1
outpaces    doc3.txt:1
over    doc1.txt:1
processes   doc5.txt:1
quick   doc1.txt:1, doc3.txt:2
sleeps  doc2.txt:1
spark   doc5.txt:1
stores  doc5.txt:1
that    doc4.txt:1
the doc1.txt:2, doc2.txt:1, doc6.txt:2
too doc4.txt:1

The Python check, 08_inverted_index.py:

OUTPUT

  Experiment 8 -- inverted index in MapReduce

    6 documents in, 26 index terms out
    map emitted 48 postings

    a slice of the index (term -> {doc: frequency}):
      dog       doc1:1, doc2:1, doc3:1, doc6:1
      quick     doc1:1, doc3:2
      big       doc4:2, doc5:2
      data      doc4:2, doc5:2
      machine   doc4:1
         'quick' appears TWICE in doc3, and the posting records
         that. Frequency is the difference between 'does this word
         occur' and 'how relevant is this document' -- boolean
         retrieval against ranked retrieval, in one number

    queries answered from the index alone:
      quick AND fox           -> ['doc1', 'doc3']
      big AND data            -> ['doc4', 'doc5']
      dog AND machine         -> no match
      dog OR machine          -> ['doc1', 'doc2', 'doc3', 'doc4', 'doc6']
         NOT ONE DOCUMENT WAS READ to answer these. That is the
         entire point of an inverted index: query cost depends on the
         number of MATCHES, not on the size of the corpus. Scanning
         6 documents is cheap; scanning 6 billion is not

    the index is not free:
      corpus         226 characters
      index terms     26
      postings        39
      ratio          6.5 postings per document
         a full-text index typically runs 20-40% of the corpus
         size, and that is BEFORE positions. Search is a space-for-time
         trade, and 'the index is bigger than I expected' is the normal
         outcome, not a mistake

    why MapReduce suits this:
      map    is per-document and EMBARRASSINGLY PARALLEL
      reduce is per-term, and every posting for a term
             arrives at the same reducer by construction
         the shuffle does the hard part -- gathering every
         mention of a word from every machine in the cluster -- and
         you never wrote a line of network code. That is the whole
         value proposition of the model

    the skew, measured:
      largest posting list : 'the' with 5
      singleton terms      : 15 of 26
         'the' is the biggest list here and would be the biggest
         on any English corpus. Real engines drop stop words or split
         hot terms across reducers, because one reducer holding 'the'
         is the job's critical path

QUERIES ANSWERED FROM THE INDEX ALONE

Query Result
quick AND fox doc1, doc3
big AND data doc4, doc5
dog AND machine no match
dog OR machine doc1, doc2, doc3, doc4, doc6

Not one document was read. Query cost depends on the number of matches, not on the size of the corpus — which is the entire point of an inverted index.

The index is not free. 226 characters of corpus → 26 terms, 39 postings. A full-text index typically runs 20–40% of the corpus size, before positions. Search is a space-for-time trade.

The skew, measured. Largest posting list: the, with 5. Singleton terms: 15 of 26. the would be the biggest list on any English corpus, and one reducer holding it is the job's critical path.

THE JAVA DETAIL THAT IS EXAMINED

The filename is in neither the key nor the value. It comes from the input split — ((FileSplit) context.getInputSplit()).getPath().getName().

And this job cannot use a combiner: the reducer's output type (a posting string) differs from its input type (a document id), and a combiner must match the reducer on both — the cluster's counters show Combine input records=0.

RESULT

6 documents in, 26 index terms out, from 48 postings — on the cluster and in the engine alike; quick is doc1:1, doc3:2. Boolean queries are answered without reading a document, and the largest posting list is the, 5 occurrences in 3 documents.

Experiment 9 — Data analysis with Pig Latin

1. Question

Perform data analysis using Pig Latin scripts.

2. Aim

Load the sales, filter, group and total them by category, order the result, join the stores map-side and flatten the tags; then walk the same dataflow one operator at a time.

3. Steps

On the cluster, 09_analysis.pig:

  1. Load the sales.
  2. Build the dataflow, a relation per step.
  3. Describe, illustrate and explain it.
  4. Store the result.
  5. Join the stores, map-side.
  6. Flatten the tags.

The Python check, 09_pig_equivalent.py:

  1. Load the sales.
  2. Filter the bulk orders.
  3. Group by category.
  4. Total each group.
  5. Order by revenue.
  6. Join the stores.
  7. Map the operators to SQL.
  8. See lazy evaluation.

THE DATAFLOW

The Python half walks the dataflow one operator at a time, which is how you debug a Pig script anyway — that is what ILLUSTRATE does.

A = LOAD 'sales'                 -- 9 rows
B = FILTER A BY qty >= 6         -- 7 rows
C = GROUP B BY category          -- 3 groups: Grocery {4}, Personal {1}, Stationery {2}
D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)
E = ORDER D BY revenue DESC      -- top category: Grocery

4. Programme

On the cluster, 09_analysis.pig:

-- Experiment 9 -- data analysis with Pig Latin
--
-- Run it: pig -x mapreduce 09_analysis.pig, with the input files on HDFS. It was run on a Hadoop 3.3.6 cluster where these labs
-- are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
-- [Changed: this said the file had never been run, as the Hadoop stack could not be
-- installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
--
-- The runnable half is 09_pig_equivalent.py, which walks the same dataflow relation by relation
--
-- run with:  pig -x mapreduce 09_analysis.pig
--       or:  pig -x local 09_analysis.pig

-- Step 1: Load the sales
sales = LOAD '/user/student/sales/sales.csv' USING PigStorage(',')
        AS (date_key:chararray, store_name:chararray, region:chararray,
            product:chararray, category:chararray,
            qty:int, list_price:double);

-- [Noted, from running it: sales.csv has a header row, and PigStorage has no
--  idea -- "Successfully read 10 records" for nine sales. The header's qty, the
--  word qty, becomes null and the FILTER below drops it without a word.
--  org.apache.pig.piggybank.storage.CSVExcelStorage can skip a header.]

-- Step 2: Build the dataflow, a relation per step
-- a relation per step. This is the point of Pig.
priced   = FOREACH sales GENERATE *, qty * list_price AS revenue;
bulk     = FILTER priced BY qty >= 6;
by_cat   = GROUP bulk BY category;
totals   = FOREACH by_cat GENERATE
               group AS category,
               COUNT(bulk)        AS orders,
               SUM(bulk.qty)      AS units,
               SUM(bulk.revenue)  AS revenue;
ranked   = ORDER totals BY revenue DESC;

-- Step 3: Describe, illustrate and explain it
DESCRIBE ranked;      -- the SCHEMA, without running anything
ILLUSTRATE ranked;    -- sample rows pushed through EVERY step -- Pig's
                      -- best feature and the one with no SQL equivalent
EXPLAIN -brief ranked;  -- logical, physical and MapReduce plans
                        -- [Changed: -brief leaves the nested plans folded; in
                        --  full they run to some 400 lines]

-- Step 4: Store the result
STORE ranked INTO '/user/student/out/by_category' USING PigStorage(',');
-- ^ NOTHING ABOVE THIS LINE HAS RUN. Pig is lazy: LOAD/FILTER/GROUP build a
--   plan, and STORE or DUMP compiles it into MapReduce jobs and submits them.

-- Step 5: Join the stores, map-side
-- a join, and the hint that decides how it is executed
stores  = LOAD '/user/student/dim/stores.csv' USING PigStorage(',')
          AS (store_name:chararray, city:chararray, region:chararray);
joined  = JOIN priced BY store_name, stores BY store_name USING 'replicated';
--                                              ^^^^^^^^^^^^^^^^^^
--   'replicated' = a MAP-SIDE join: the small relation is loaded into memory
--   on every mapper, so there is NO SHUFFLE. It fails with an OOM if the
--   right-hand relation does not fit. The default is a reduce-side join.
--   Other strategies: 'skewed' (for one hot key), 'merge' (both sorted).
-- [Corrected: the field was called store, and STORE is a Pig keyword in any
--  case. LOAD ... AS accepted it, but JOIN ... BY store, stores BY store
--  failed to parse -- and Pig parses the whole script before it runs any
--  STORE, so the error lost the by_category output above as well.]
DUMP joined;
-- [Changed: DUMP joined is added. Nothing used joined, and Pig is lazy, so
--  the join was never run.]

-- Step 6: Flatten the tags
-- FLATTEN, which has no clean SQL equivalent
tags    = LOAD '/user/student/tags.csv' AS (product:chararray, taglist:chararray);
split_t = FOREACH tags GENERATE product,
              FLATTEN(TOKENIZE(taglist, ';')) AS tag;
DUMP split_t;   -- one row per (product, tag) pair, from one row per product

The Python check, 09_pig_equivalent.py:

"""Experiment 9 -- data analysis with Pig Latin scripts.

`09_analysis.pig` carries the real Pig Latin, and runs on Pig itself (the lab
page shows it). What runs here is the same dataflow, one operator at a time, so the
INTERMEDIATE relations in the notes are real -- and stepping through them is
exactly how you debug a Pig script anyway (that is what ILLUSTRATE does).

Pig's value over Hive is that it is a DATAFLOW language: you name every
intermediate relation, so a 12-step transformation reads top to bottom instead
of nesting twelve sub-queries.
"""
import fixtures as f

SALES = f.SALES_DF


def show(name, rel, cols, limit=4):
    print(f"\n    {name} -- {len(rel)} rows")
    head = rel[cols].head(limit)
    widths = [max(len(str(c)), int(rel[c].astype(str).str.len().max())) + 3
              for c in cols]
    print("      " + "".join(f"{c:>{w}}" for c, w in zip(cols, widths)))
    for _, r in head.iterrows():
        print("      " + "".join(
            f"{r[c]:>{w},.0f}" if isinstance(r[c], float) else f"{str(r[c]):>{w}}"
            for c, w in zip(cols, widths)))
    if len(rel) > limit:
        print(f"      ... {len(rel) - limit} more")


def main():
    print("  Experiment 9 -- the Pig Latin dataflow, one operator at a time")

    # Step 1: Load the sales
    # A = LOAD
    A = SALES.copy()
    show("A = LOAD 'sales'", A, ["store", "product", "qty", "revenue"])

    # Step 2: Filter the bulk orders
    # B = FILTER
    B = A[A["qty"] >= 6]
    show("B = FILTER A BY qty >= 6", B, ["store", "product", "qty", "revenue"])
    assert len(B) == 7, 'seven of nine orders are 6 units or more'

    # Step 3: Group by category
    # C = GROUP
    C = B.groupby("category")
    print(f"\n    C = GROUP B BY category -- {C.ngroups} groups")
    for name, grp in C:
        print(f"      ({name}, {{{len(grp)} tuples}})")
    print("""         GROUP in Pig produces a BAG per key, not an aggregate.
         The bag is the value, and FOREACH ... GENERATE is what turns
         it into numbers. Hive fuses the two; Pig keeps them apart,
         which is why Pig can do things to a group that SQL cannot
         express without a window function""")

    # Step 4: Total each group
    # D = FOREACH ... GENERATE
    D = (C.agg(units=("qty", "sum"), revenue=("revenue", "sum"),
               orders=("order_id", "count") if "order_id" in B else ("qty", "size"))
         .reset_index())
    print(f"\n    D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)")
    print(f"      {'category':<12}{'units':>8}{'revenue':>12}{'orders':>8}")
    for _, r in D.iterrows():
        print(f"      {r['category']:<12}{r['units']:>8.0f}"
              f"{r['revenue']:>12,.0f}{r['orders']:>8.0f}")
    assert D["revenue"].sum() == B["revenue"].sum()

    # Step 5: Order by revenue
    # E = ORDER
    E = D.sort_values("revenue", ascending=False)
    print(f"\n    E = ORDER D BY revenue DESC")
    print(f"      top category: {E.iloc[0]['category']} "
          f"at {E.iloc[0]['revenue']:,.0f}")
    assert E.iloc[0]["category"] == "Grocery"

    # Step 6: Join the stores
    # F = JOIN
    print("\n    F = JOIN A BY store_key, stores BY store_key")
    joined = A.groupby(["region", "store"], as_index=False)["revenue"].sum()
    print(f"      {'region':<8}{'store':<14}{'revenue':>12}")
    for _, r in joined.sort_values("revenue", ascending=False).iterrows():
        print(f"      {r['region']:<8}{r['store']:<14}{r['revenue']:>12,.0f}")
    assert joined["revenue"].sum() == f.total_revenue()

    # Step 7: Map the operators to SQL
    print("\n    the operators, and their SQL equivalents:")
    print(f"      {'Pig Latin':<26}{'SQL'}")
    for pig, sql in (
            ("LOAD / STORE", "no equivalent -- SQL assumes a table exists"),
            ("FILTER", "WHERE"),
            ("FOREACH .. GENERATE", "SELECT"),
            ("GROUP", "GROUP BY, but the bag is kept"),
            ("JOIN", "JOIN"),
            ("ORDER", "ORDER BY"),
            ("DISTINCT", "DISTINCT"),
            ("FLATTEN", "UNNEST / LATERAL VIEW explode"),
            ("ILLUSTRATE", "no equivalent -- sample data through the plan")):
        print(f"      {pig:<26}{sql}")

    print("""
         two operators have no SQL equivalent, and they are the
         reason to reach for Pig: LOAD, which lets a script read a
         semi-structured file with no schema declared in advance,
         and ILLUSTRATE, which pushes a few representative rows
         through every step of the plan so you can see where a
         12-stage pipeline went wrong""")

    # Step 8: See lazy evaluation
    print("\n    lazy evaluation, which surprises everyone:")
    print("      nothing runs until STORE or DUMP.")
    print("      A = LOAD ...;   B = FILTER ...;   C = GROUP ...;")
    print("      -- no job has been submitted yet")
    print("      STORE C INTO 'out';  -- NOW Pig compiles and runs it")
    print("""         because Pig sees the whole dataflow before executing, it
         can merge the FILTER into the LOAD and fuse consecutive
         FOREACHes into one MapReduce job. Writing the steps
         separately costs nothing -- which is the entire argument
         against nesting sub-queries to avoid 'extra passes'""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 09_analysis.pig:

OUTPUT

$ hdfs dfs -mkdir -p /user/student/sales /user/student/dim
$ hdfs dfs -put sales.csv /user/student/sales/
$ hdfs dfs -put stores.csv /user/student/dim/
$ hdfs dfs -put tags.tsv /user/student/tags.csv
$ pig -x mapreduce 09_analysis.pig 2>pig.log
ranked: {category: chararray,orders: long,units: long,revenue: double}
(D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0)
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| sales     | date_key:chararray    | store_name:chararray    | region:chararray    | product:chararray    | category:chararray    | qty:int    | list_price:double    |
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|           | D3                    | Hyderabad               | North               | Rice 5kg             | Grocery               | 4          | 280.0                |
|           | D3                    | Guntur                  | South               | Tea 500g             | Grocery               | 12         | 210.0                |
|           | D2                    | Vijayawada              | South               | Rice 5kg             | Grocery               | 6          | 280.0                |
|           | D4                    | Hyderabad               | North               | Notebook             | Stationery            | 15         | 40.0                 |
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| priced     | date_key:chararray    | store_name:chararray    | region:chararray    | product:chararray    | category:chararray    | qty:int    | list_price:double    | revenue:double     |
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|            | D3                    | Hyderabad               | North               | Rice 5kg             | Grocery               | 4          | 280.0                | 1120.0             |
|            | D3                    | Guntur                  | South               | Tea 500g             | Grocery               | 12         | 210.0                | 2520.0             |
|            | D2                    | Vijayawada              | South               | Rice 5kg             | Grocery               | 6          | 280.0                | 1680.0             |
|            | D4                    | Hyderabad               | North               | Notebook             | Stationery            | 15         | 40.0                 | 600.0              |
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| bulk     | date_key:chararray    | store_name:chararray    | region:chararray    | product:chararray    | category:chararray    | qty:int    | list_price:double    | revenue:double     |
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|          | D3                    | Guntur                  | South               | Tea 500g             | Grocery               | 12         | 210.0                | 2520.0             |
|          | D2                    | Vijayawada              | South               | Rice 5kg             | Grocery               | 6          | 280.0                | 1680.0             |
|          | D4                    | Hyderabad               | North               | Notebook             | Stationery            | 15         | 40.0                 | 600.0              |
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| by_cat     | group:chararray    | bulk:bag{:tuple(date_key:chararray,store_name:chararray,region:chararray,product:chararray,category:chararray,qty:int,list_price:double,revenue:double)}                                  |
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|            | Grocery            | {}                                                                                                                                                                                        |
|            | Grocery            | {}                                                                                                                                                                                        |
|            | Stationery         | {}                                                                                                                                                                                        |
|            | Stationery         | {}                                                                                                                                                                                        |
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------
| totals     | category:chararray    | orders:long     | units:long     | revenue:double     |
----------------------------------------------------------------------------------------------
|            | Grocery               | 2               | 18             | 4200.0             |
|            | Stationery            | 1               | 15             | 600.0              |
----------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------
| ranked     | category:chararray    | orders:long     | units:long     | revenue:double     |
----------------------------------------------------------------------------------------------
|            | Grocery               | 2               | 18             | 4200.0             |
|            | Stationery            | 1               | 15             | 600.0              |
----------------------------------------------------------------------------------------------

#-----------------------------------------------
# New Logical Plan:
#-----------------------------------------------
ranked: (Name: LOStore Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
|
|---ranked: (Name: LOSort Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
    |   |
    |   revenue:(Name: Project Type: double Uid: 186 Input: 0 Column: 3)
    |
    |---totals: (Name: LOForEach Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
        |   |
        |   (Name: LOGenerate[false,false,false,false] Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)ColumnPrune:OutputUids=[180, 183, 152, 186]ColumnPrune:InputUids=[178, 152]
        |   |   |
        |   |   group:(Name: Project Type: chararray Uid: 152 Input: 0 Column: (*))
        |   |   |
        |   |   (Name: UserFunc(org.apache.pig.builtin.COUNT) Type: long Uid: 180)
        |   |   |
        |   |   |---bulk:(Name: Project Type: bag Uid: 178 Input: 1 Column: (*))
        |   |   |
        |   |   (Name: UserFunc(org.apache.pig.builtin.LongSum) Type: long Uid: 183)
        |   |   |
        |   |   |---(Name: Dereference Type: bag Uid: 182 Column:[5])
        |   |       |
        |   |       |---bulk:(Name: Project Type: bag Uid: 178 Input: 2 Column: (*))
        |   |   |
        |   |   (Name: UserFunc(org.apache.pig.builtin.DoubleSum) Type: double Uid: 186)
        |   |   |
        |   |   |---(Name: Dereference Type: bag Uid: 185 Column:[7])
        |   |       |
        |   |       |---bulk:(Name: Project Type: bag Uid: 178 Input: 3 Column: (*))
        |   |
        |   |---(Name: LOInnerLoad[0] Schema: group#152:chararray)
        |   |
        |   |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
        |   |
        |   |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
        |   |
        |   |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
        |
        |---by_cat: (Name: LOCogroup Schema: group#152:chararray,priced#178:bag{#199:tuple(date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)})
            |   |
            |   category:(Name: Project Type: chararray Uid: 152 Input: 0 Column: 4)
            |
            |---priced: (Name: LOForEach Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
                |   |
                |   (Name: LOGenerate[false,false,false,false,false,false,false,false] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)ColumnPrune:OutputUids=[148, 149, 150, 151, 152, 153, 154, 175]ColumnPrune:InputUids=[148, 149, 150, 151, 152, 153, 154]
                |   |   |
                |   |   date_key:(Name: Project Type: chararray Uid: 148 Input: 2 Column: (*))
                |   |   |
                |   |   store_name:(Name: Project Type: chararray Uid: 149 Input: 3 Column: (*))
                |   |   |
                |   |   region:(Name: Project Type: chararray Uid: 150 Input: 4 Column: (*))
                |   |   |
                |   |   product:(Name: Project Type: chararray Uid: 151 Input: 5 Column: (*))
                |   |   |
                |   |   category:(Name: Project Type: chararray Uid: 152 Input: 6 Column: (*))
                |   |   |
                |   |   qty:(Name: Project Type: int Uid: 153 Input: 7 Column: (*))
                |   |   |
                |   |   list_price:(Name: Project Type: double Uid: 154 Input: 8 Column: (*))
                |   |   |
                |   |   (Name: Multiply Type: double Uid: 175)
                |   |   |
                |   |   |---(Name: Cast Type: double Uid: 153)
                |   |   |   |
                |   |   |   |---qty:(Name: Project Type: int Uid: 153 Input: 0 Column: (*))
                |   |   |
                |   |   |---list_price:(Name: Project Type: double Uid: 154 Input: 1 Column: (*))
                |   |
                |   |---(Name: LOInnerLoad[5] Schema: qty#153:int)
                |   |
                |   |---(Name: LOInnerLoad[6] Schema: list_price#154:double)
                |   |
                |   |---(Name: LOInnerLoad[0] Schema: date_key#148:chararray)
                |   |
                |   |---(Name: LOInnerLoad[1] Schema: store_name#149:chararray)
                |   |
                |   |---(Name: LOInnerLoad[2] Schema: region#150:chararray)
                |   |
                |   |---(Name: LOInnerLoad[3] Schema: product#151:chararray)
                |   |
                |   |---(Name: LOInnerLoad[4] Schema: category#152:chararray)
                |   |
                |   |---(Name: LOInnerLoad[5] Schema: qty#153:int)
                |   |
                |   |---(Name: LOInnerLoad[6] Schema: list_price#154:double)
                |
                |---bulk: (Name: LOFilter Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)
                    |   |
                    |   (Name: GreaterThanEqual Type: boolean Uid: 177)
                    |   |
                    |   |---qty:(Name: Project Type: int Uid: 153 Input: 0 Column: 5)
                    |   |
                    |   |---(Name: Constant Type: int Uid: 176)
                    |
                    |---sales: (Name: LOForEach Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)
                        |   |
                        |   (Name: LOGenerate[false,false,false,false,false,false,false] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)ColumnPrune:OutputUids=[148, 149, 150, 151, 152, 153, 154]ColumnPrune:InputUids=[148, 149, 150, 151, 152, 153, 154]
                        |   |   |
                        |   |   (Name: Cast Type: chararray Uid: 148)
                        |   |   |
                        |   |   |---date_key:(Name: Project Type: bytearray Uid: 148 Input: 0 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: chararray Uid: 149)
                        |   |   |
                        |   |   |---store_name:(Name: Project Type: bytearray Uid: 149 Input: 1 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: chararray Uid: 150)
                        |   |   |
                        |   |   |---region:(Name: Project Type: bytearray Uid: 150 Input: 2 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: chararray Uid: 151)
                        |   |   |
                        |   |   |---product:(Name: Project Type: bytearray Uid: 151 Input: 3 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: chararray Uid: 152)
                        |   |   |
                        |   |   |---category:(Name: Project Type: bytearray Uid: 152 Input: 4 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: int Uid: 153)
                        |   |   |
                        |   |   |---qty:(Name: Project Type: bytearray Uid: 153 Input: 5 Column: (*))
                        |   |   |
                        |   |   (Name: Cast Type: double Uid: 154)
                        |   |   |
                        |   |   |---list_price:(Name: Project Type: bytearray Uid: 154 Input: 6 Column: (*))
                        |   |
                        |   |---(Name: LOInnerLoad[0] Schema: date_key#148:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[1] Schema: store_name#149:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[2] Schema: region#150:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[3] Schema: product#151:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[4] Schema: category#152:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[5] Schema: qty#153:bytearray)
                        |   |
                        |   |---(Name: LOInnerLoad[6] Schema: list_price#154:bytearray)
                        |
                        |---sales: (Name: LOLoad Schema: date_key#148:bytearray,store_name#149:bytearray,region#150:bytearray,product#151:bytearray,category#152:bytearray,qty#153:bytearray,list_price#154:bytearray)RequiredFields:null
#-----------------------------------------------
# Physical Plan:
#-----------------------------------------------
ranked: Store(fakefile:org.apache.pig.builtin.PigStorage) - scope-804
|
|---ranked: POSort[bag]() - scope-803
    |
    |---totals: New For Each(false,false,false,false)[bag] - scope-801
        |
        |---by_cat: Package(Packager)[tuple]{chararray} - scope-785
            |
            |---by_cat: Global Rearrange[tuple] - scope-784
                |
                |---by_cat: Local Rearrange[tuple]{chararray}(false) - scope-786
                    |
                    |---priced: New For Each(false,false,false,false,false,false,false,false)[bag] - scope-783
                        |
                        |---bulk: Filter[bag] - scope-759
                            |
                            |---sales: New For Each(false,false,false,false,false,false,false)[bag] - scope-758
                                |
                                |---sales: Load(/user/student/sales/sales.csv:PigStorage(',')) - scope-736

#--------------------------------------------------
# Map Reduce Plan
#--------------------------------------------------
MapReduce node scope-805
Map Plan
by_cat: Local Rearrange[tuple]{chararray}(false) - scope-851
|
|---totals: New For Each(false,false,false,false)[bag] - scope-828
    |
    |---Pre Combiner Local Rearrange[tuple]{Unknown} - scope-854
        |
        |---priced: New For Each(false,false,false,false,false,false,false,false)[bag] - scope-783
            |
            |---bulk: Filter[bag] - scope-759
                |
                |---sales: New For Each(false,false,false,false,false,false,false)[bag] - scope-758
                    |
                    |---sales: Load(/user/student/sales/sales.csv:PigStorage(',')) - scope-736--------
Combine Plan
by_cat: Local Rearrange[tuple]{chararray}(false) - scope-855
|
|---totals: New For Each(false,false,false,false)[bag] - scope-838
    |
    |---by_cat: Package(CombinerPackager)[tuple]{chararray} - scope-850--------
Reduce Plan
Store(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.io.InterStorage) - scope-806
|
|---totals: New For Each(false,false,false,false)[bag] - scope-801
    |
    |---by_cat: Package(CombinerPackager)[tuple]{chararray} - scope-785--------
Global sort: false
----------------

MapReduce node scope-808
Map Plan
ranked: Local Rearrange[tuple]{tuple}(false) - scope-812
|
|---New For Each(false)[tuple] - scope-810
    |
    |---Load(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.builtin.RandomSampleLoader('org.apache.pig.impl.io.InterStorage','100')) - scope-807--------
Reduce Plan
Store(hdfs://localhost:9000/tmp/temp1251966677/tmp2048799314:org.apache.pig.impl.io.InterStorage) - scope-821
|
|---New For Each(false)[tuple] - scope-820
    |
    |---New For Each(false,false)[tuple] - scope-817
        |
        |---Package(Packager)[tuple]{chararray} - scope-813--------
Global sort: false
Secondary sort: true
----------------

MapReduce node scope-823
Map Plan
ranked: Local Rearrange[tuple]{double}(false) - scope-824
|
|---Load(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.io.InterStorage) - scope-822--------
Reduce Plan
ranked: Store(fakefile:org.apache.pig.builtin.PigStorage) - scope-804
|
|---New For Each(true)[tuple] - scope-827
    |
    |---Package(LitePackager)[tuple]{double} - scope-825--------
Global sort: true
Quantile file: hdfs://localhost:9000/tmp/temp1251966677/tmp2048799314
----------------

(D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0,2800.0,Vijayawada,Vijayawada,South)
(D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0,700.0,Vijayawada,Vijayawada,South)
(D1,Guntur,South,Tea 500g,Grocery,8,210.0,1680.0,Guntur,Guntur,South)
(D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0,1680.0,Vijayawada,Vijayawada,South)
(D2,Hyderabad,North,Notebook,Stationery,20,40.0,800.0,Hyderabad,Hyderabad,North)
(D3,Guntur,South,Tea 500g,Grocery,12,210.0,2520.0,Guntur,Guntur,South)
(D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0,1120.0,Hyderabad,Hyderabad,North)
(D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0,980.0,Vijayawada,Vijayawada,South)
(D4,Hyderabad,North,Notebook,Stationery,15,40.0,600.0,Hyderabad,Hyderabad,North)
(Rice 5kg,staple)
(Rice 5kg,grain)
(Rice 5kg,bulk)
(Tea 500g,beverage)
(Tea 500g,daily)
(Shampoo 200ml,personal)
(Shampoo 200ml,care)
(Notebook,paper)
(Notebook,school)
$ grep -E 'ERROR|Success|Failed' pig.log | sed -E 's/^[0-9-]+ [0-9:,]+ //' | head -12
Success!
Successfully read 10 records (845 bytes) from: "/user/student/sales/sales.csv"
Successfully stored 3 records (62 bytes) in: "/user/student/out/by_category"
[main] INFO  org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.MapReduceLauncher - Success!
Success!
Successfully read 3 records (450 bytes) from: "/user/student/dim/stores.csv"
Successfully read 10 records (845 bytes) from: "/user/student/sales/sales.csv"
Successfully stored 9 records (939 bytes) in: "hdfs://localhost:9000/tmp/temp1251966677/tmp-1038105655"
[main] INFO  org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.MapReduceLauncher - Success!
Success!
Successfully read 4 records (471 bytes) from: "/user/student/tags.csv"
Successfully stored 9 records (223 bytes) in: "hdfs://localhost:9000/tmp/temp1251966677/tmp1005279582"
$ hdfs dfs -cat /user/student/out/by_category/part-r-00000
Grocery,4,36,8680.0
Stationery,2,35,1400.0
Personal,1,7,980.0

The Python check, 09_pig_equivalent.py:

OUTPUT

  Experiment 9 -- the Pig Latin dataflow, one operator at a time

    A = LOAD 'sales' -- 9 rows
              store         product   qty   revenue
         Vijayawada        Rice 5kg    10     2,800
         Vijayawada   Shampoo 200ml     5       700
             Guntur        Tea 500g     8     1,680
         Vijayawada        Rice 5kg     6     1,680
      ... 5 more

    B = FILTER A BY qty >= 6 -- 7 rows
              store         product   qty   revenue
         Vijayawada        Rice 5kg    10     2,800
             Guntur        Tea 500g     8     1,680
         Vijayawada        Rice 5kg     6     1,680
          Hyderabad        Notebook    20       800
      ... 3 more

    C = GROUP B BY category -- 3 groups
      (Grocery, {4 tuples})
      (Personal, {1 tuples})
      (Stationery, {2 tuples})
         GROUP in Pig produces a BAG per key, not an aggregate.
         The bag is the value, and FOREACH ... GENERATE is what turns
         it into numbers. Hive fuses the two; Pig keeps them apart,
         which is why Pig can do things to a group that SQL cannot
         express without a window function

    D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)
      category       units     revenue  orders
      Grocery           36       8,680       4
      Personal           7         980       1
      Stationery        35       1,400       2

    E = ORDER D BY revenue DESC
      top category: Grocery at 8,680

    F = JOIN A BY store_key, stores BY store_key
      region  store              revenue
      South   Vijayawada           6,160
      South   Guntur               4,200
      North   Hyderabad            2,520

    the operators, and their SQL equivalents:
      Pig Latin                 SQL
      LOAD / STORE              no equivalent -- SQL assumes a table exists
      FILTER                    WHERE
      FOREACH .. GENERATE       SELECT
      GROUP                     GROUP BY, but the bag is kept
      JOIN                      JOIN
      ORDER                     ORDER BY
      DISTINCT                  DISTINCT
      FLATTEN                   UNNEST / LATERAL VIEW explode
      ILLUSTRATE                no equivalent -- sample data through the plan

         two operators have no SQL equivalent, and they are the
         reason to reach for Pig: LOAD, which lets a script read a
         semi-structured file with no schema declared in advance,
         and ILLUSTRATE, which pushes a few representative rows
         through every step of the plan so you can see where a
         12-stage pipeline went wrong

    lazy evaluation, which surprises everyone:
      nothing runs until STORE or DUMP.
      A = LOAD ...;   B = FILTER ...;   C = GROUP ...;
      -- no job has been submitted yet
      STORE C INTO 'out';  -- NOW Pig compiles and runs it
         because Pig sees the whole dataflow before executing, it
         can merge the FILTER into the LOAD and fuse consecutive
         FOREACHes into one MapReduce job. Writing the steps
         separately costs nothing -- which is the entire argument
         against nesting sub-queries to avoid 'extra passes'

GROUP PRODUCES A BAG, NOT AN AGGREGATE

(Grocery, {4 tuples}) — the bag is the value, and FOREACH … GENERATE turns it into numbers. Hive fuses the two; Pig keeps them apart, which is why Pig can do things to a group that SQL cannot express without a window function.

The two operators with no SQL equivalent. LOAD reads a semi-structured file with no schema declared in advance — SQL assumes a table exists. ILLUSTRATE pushes sample rows through every step of the plan. Those two are the reason to reach for Pig on ETL.

Lazy evaluation. Nothing runs until STORE or DUMP. Pig sees the whole dataflow first, so it merges the FILTER into the LOAD and fuses consecutive FOREACHes into one job — the EXPLAIN above shows the plan it made. Writing the steps separately costs nothing — which is the whole argument against nesting sub-queries "to avoid extra passes".

The join hint that matters. USING 'replicated' is a map-side join: the small relation loads into every mapper's memory and there is no shuffle. It dies with an OutOfMemoryError if the relation does not fit — which is why broadcast joins have a size threshold.

RESULT

Pig, on the cluster, gives Grocery 4 orders, 36 units, ₹8,680; Stationery 2, 35, ₹1,400; Personal 1, 7, ₹980 — the bulk orders by category, revenue first. store is a Pig keyword and cannot name a field.

Experiment 10 — Hive: tables, partitions and buckets

1. Question

Execute Hive queries for structured data analysis, with tables and partitions.

2. Aim

Lay a table over the sales file, load a partitioned and bucketed table from it, aggregate by region, prune a partition and rank with a window function; then check each figure in DuckDB.

3. Steps

On the cluster, 10_hive.hql:

  1. Create the database.
  2. Lay an external table over the file.
  3. Create the partitioned, bucketed table.
  4. Load it with dynamic partitions.
  5. Aggregate by region.
  6. Prune a partition.
  7. Rank with a window function.
  8. Tune, and compute statistics.

The Python check, 10_hive_duckdb.py:

  1. Aggregate by region.
  2. Partition by quarter.
  3. Bucket by store.
  4. Compare managed with external tables.
  5. Join, as Hive does.
  6. See what Hive is not.

THE CROSS-COURSE CHECK

Region Revenue Profit Margin
South 10,360 2,760 26.64%
North 2,520 765 30.36%

South = ₹10,360 is the same number Business Intelligence Tools' DAX CALCULATE measure produced, and the same one Spark produces in experiment 17. Hive printed it from the cluster; DuckDB, running the same query text, asserts it. Three engines, two languages, one dataset — asserted, so drift fails the suite.

4. Programme

On the cluster, 10_hive.hql:

-- Experiment 10 -- Hive queries for structured data analysis -- tables and partitions
--
-- Run it: hive -f 10_hive.hql, with sales.csv in /user/student/sales. It was run on a Hadoop 3.3.6 cluster where these labs
-- are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
-- [Changed: this said the file had never been run, as the Hadoop stack could not be
-- installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
--
-- The runnable half is 10_hive_duckdb.py, which runs the same SQL through DuckDB
--
-- run with:  hive -f 10_hive.hql       or  beeline -u jdbc:hive2://localhost:10000

-- Step 1: Create the database
CREATE DATABASE IF NOT EXISTS retail;
USE retail;

-- Step 2: Lay an external table over the file
-- an EXTERNAL table over data you did not produce: DROP TABLE will not
-- delete the files. Use this for anything you cannot recreate.
CREATE EXTERNAL TABLE IF NOT EXISTS sales_raw (
    date_key   STRING,
    store      STRING,
    region     STRING,
    product    STRING,
    category   STRING,
    qty        INT,
    list_price DOUBLE
)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
STORED AS TEXTFILE
LOCATION '/user/student/sales/'
TBLPROPERTIES ('skip.header.line.count'='1');

-- Step 3: Create the partitioned, bucketed table
-- the table you actually query: PARTITIONED and columnar
CREATE TABLE IF NOT EXISTS sales (
    date_key   STRING,
    store      STRING,
    region     STRING,
    product    STRING,
    category   STRING,
    qty        INT,
    list_price DOUBLE,
    revenue    DOUBLE
)
PARTITIONED BY (quarter STRING)
CLUSTERED BY (store) INTO 3 BUCKETS
STORED AS PARQUET;

-- Step 4: Load it with dynamic partitions
-- dynamic partitioning: Hive reads the LAST select column as the partition
SET hive.exec.dynamic.partition = true;
SET hive.exec.dynamic.partition.mode = nonstrict;

INSERT OVERWRITE TABLE sales PARTITION (quarter)
SELECT date_key, store, region, product, category, qty, list_price,
       qty * list_price AS revenue,
       CASE WHEN date_key IN ('D1','D2') THEN 'Q1' ELSE 'Q2' END AS quarter
FROM sales_raw;

SHOW PARTITIONS sales;
DESCRIBE FORMATTED sales;

-- Step 5: Aggregate by region
-- the aggregate verified against Course 11's DAX and experiment 10's DuckDB
SELECT region, SUM(revenue) AS revenue, SUM(qty) AS units
FROM sales
GROUP BY region
ORDER BY revenue DESC;
--   expected: South 10360, North 2520

-- Step 6: Prune a partition
-- PARTITION PRUNING: read one directory, not the table
EXPLAIN DEPENDENCY SELECT SUM(revenue) FROM sales WHERE quarter = 'Q2';
--   input_partitions lists quarter=Q2 alone: one directory is read. If the
--   filter is on a non-partition column it lists every partition, and no
--   amount of indexing will save it -- Hive dropped indexes in version 3.
-- [Corrected: this was a plain EXPLAIN, with the advice to look for
--  "partition values:[Q2]" in the plan. Hive 3's plan does not print that
--  line; its TableScan says 4 rows (Q2's four of nine) but names no partition.
--  EXPLAIN DEPENDENCY names the partitions a query reads.]

SELECT category, product, SUM(qty) AS units, SUM(revenue) AS revenue
FROM sales
WHERE quarter = 'Q2'          -- prunes DIRECTORIES, before the job starts
GROUP BY category, product
HAVING SUM(revenue) > 1000    -- filters GROUPS, in the reducer
ORDER BY revenue DESC;

-- Step 7: Rank with a window function
-- a window function, which is where Hive stops looking like MapReduce
SELECT region, product, revenue,
       RANK() OVER (PARTITION BY region ORDER BY revenue DESC) AS rnk
FROM sales;

-- Step 8: Tune, and compute statistics
-- the settings worth knowing
-- SET hive.execution.engine = tez;      -- mr is deprecated and slow
-- [Corrected: this line ran. Tez is a separate install, and where it is not
--  installed the next query dies with NoClassDefFoundError:
--  org/apache/tez/runtime/api/Event -- and the CLI then hangs instead of
--  exiting. Set it only where Tez is installed; this runs on MapReduce.]
SET hive.vectorized.execution.enabled = true;
SET hive.cbo.enable = true;              -- cost-based optimiser, needs stats
ANALYZE TABLE sales PARTITION(quarter) COMPUTE STATISTICS FOR COLUMNS;

The Python check, 10_hive_duckdb.py:

"""Experiment 10 -- Hive queries for structured data analysis: tables,
partitions and the queries that go with them.

`10_hive.hql` carries the HiveQL you submit, and runs on Hive itself (the lab
page shows it). This runs the same questions through DuckDB, which speaks close
enough to ANSI SQL that the SAME query text answers them -- so the figures are
asserted here in seconds, and Hive's answers must agree with them.

The data is Course 11's star schema, imported rather than copied, so a Hive
aggregate here and a DAX measure there are computed from the same nine rows.
"""
import duckdb

import fixtures as f


def q(con, sql):
    return con.execute(sql).fetchall()


def main():
    print("  Experiment 10 -- Hive-style SQL over the star schema")

    con = duckdb.connect()
    con.register("sales", f.SALES_DF)

    print(f"\n    {len(f.SALES_DF)} fact rows, "
          f"total revenue {f.total_revenue():,.0f}")

    # Step 1: Aggregate by region
    rows = q(con, """
        SELECT region, SUM(revenue) AS revenue, SUM(profit) AS profit
        FROM sales GROUP BY region ORDER BY revenue DESC
    """)
    print(f"\n    {'region':<10}{'revenue':>12}{'profit':>10}{'margin':>9}")
    for region, rev, prof in rows:
        print(f"    {region:<10}{rev:>12,.0f}{prof:>10,.0f}{100 * prof / rev:>8.2f}%")
    got = dict((r[0], r[1]) for r in rows)
    assert got["South"] == 10360.0, "must match Course 11's CALCULATE figure"
    assert got["North"] == 2520.0
    assert sum(got.values()) == f.total_revenue()
    print("""         South = 10,360 is the SAME number Course 11's DAX
         CALCULATE measure produced. Two engines, two languages, one
         dataset -- if they ever disagree the suite fails, which is
         what makes the cross-check worth having""")

    # Step 2: Partition by quarter
    print("\n    partitioning by quarter -- what Hive actually does:")
    parts = q(con, """
        SELECT quarter, COUNT(*) AS rows, SUM(revenue) AS revenue
        FROM sales GROUP BY quarter ORDER BY quarter
    """)
    total_rows = sum(p[1] for p in parts)
    print(f"      {'partition':<14}{'rows':>6}{'revenue':>12}{'scanned for Q2':>16}")
    for quarter, n, rev in parts:
        print(f"      quarter={quarter:<7}{n:>6}{rev:>12,.0f}"
              f"{(n if quarter == 'Q2' else 0):>16}")
    q2_rows = next(n for quarter, n, _ in parts if quarter == "Q2")
    print(f"      {'TOTAL':<14}{total_rows:>6}{'':>12}{q2_rows:>16}")
    assert total_rows == 9 and q2_rows == 4
    print(f"""         a partitioned table stores each quarter in its own HDFS
         DIRECTORY, so 'WHERE quarter = ''Q2''' reads {q2_rows} rows instead
         of {total_rows} -- partition PRUNING, decided before a single byte is
         read. The partition column is a directory name, not a column
         in the data files, which is why it costs no storage""")

    print("""
      the trap: partition on something with FEW distinct values.
      Partitioning by date_key here would make 4 directories for 9
      rows -- the small-files problem from experiment 4, created on
      purpose. Partition by quarter or month; BUCKET by customer_id""")

    # Step 3: Bucket by store
    print("\n    bucketing (CLUSTERED BY store INTO 3 BUCKETS):")
    buckets = {}
    for store in sorted(f.SALES_DF["store"].unique()):
        h = sum(ord(c) for c in store) % 3
        buckets.setdefault(h, []).append(store)
    for b in range(3):
        print(f"      bucket {b}: {buckets.get(b, []) or '(empty)'}")
    assert 1 not in buckets, "bucket 1 draws nothing from three store names"
    print("""         BUCKET 1 IS EMPTY, with three stores over three buckets.
         Hashing does not distribute small key sets evenly, and an
         empty bucket is still a file the job opens.
         Buckets are FILES inside a partition, assigned by a hash
         of the column. Two tables bucketed the same way on the same
         column can be joined bucket-to-bucket with no shuffle at all
         -- a sort-merge bucket join, and the reason bucketing exists""")

    # Step 4: Compare managed with external tables
    print("\n    managed against external tables:")
    print(f"      {'':<12}{'data lives':<26}{'DROP TABLE deletes'}")
    print(f"      {'MANAGED':<12}{'/user/hive/warehouse':<26}{'the DATA too'}")
    print(f"      {'EXTERNAL':<12}{'wherever you point it':<26}{'only the metadata'}")
    print("""         use EXTERNAL for data you did not produce and cannot
         recreate. A DROP TABLE on a managed table over the company's
         only copy of a dataset is the classic Hive accident""")

    # Step 5: Join, as Hive does
    rows = q(con, """
        SELECT category, product, SUM(qty) AS units, SUM(revenue) AS revenue
        FROM sales GROUP BY category, product
        HAVING SUM(revenue) > 1000
        ORDER BY revenue DESC
    """)
    print(f"\n    products above 1,000 revenue:")
    print(f"      {'category':<12}{'product':<16}{'units':>6}{'revenue':>10}")
    for cat, prod, units, rev in rows:
        print(f"      {cat:<12}{prod:<16}{units:>6.0f}{rev:>10,.0f}")
    assert len(rows) == 4, "four products clear 1,000 -- Notebook only just"
    grocery = sum(r[3] for r in rows if r[0] == "Grocery")
    assert grocery == 9800.0
    print("""         HAVING filters GROUPS, WHERE filters ROWS -- and in Hive
         that distinction is a job-plan difference, not a syntax
         nicety: a WHERE on a partition column prunes directories
         before the job starts, a HAVING cannot""")

    # Step 6: See what Hive is not
    print("\n    what Hive is NOT:")
    print(f"      {'expectation':<30}{'reality'}")
    for exp, real in (
            ("row-level UPDATE/DELETE", "only with ACID tables + ORC + buckets"),
            ("sub-second queries", "seconds to minutes -- it plans a JOB"),
            ("indexes", "removed in Hive 3; use partitions and ORC/Parquet"),
            ("a running server holding data", "metadata only; data is files in HDFS"),
            ("enforced constraints", "declarative only; NOT enforced")):
        print(f"      {exp:<30}{real}")
    print("""         Hive is a COMPILER: HiveQL in, a MapReduce/Tez/Spark job
         out. Everything surprising about it follows from that one
         sentence, and it is the right answer to 'compare Hive with an
         RDBMS'""")

    con.close()


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 10_hive.hql:

OUTPUT

$ hdfs dfs -mkdir -p /user/student/sales /user/hive/warehouse /tmp/hive
$ hdfs dfs -chmod 733 /tmp/hive && hdfs dfs -chmod g+w /user/hive/warehouse
$ hdfs dfs -put sales.csv /user/student/sales/
$ schematool -dbType derby -initSchema 2>&1 | grep -E 'schemaTool completed'
schemaTool completed
$ hive -f 10_hive.hql 2>hive.log | grep -v '^WARN: '
quarter=Q1
quarter=Q2
# col_name              data_type               comment
date_key                string
store                   string
region                  string
product                 string
category                string
qty                     int
list_price              double
revenue                 double

# Partition Information
# col_name              data_type               comment
quarter                 string

# Detailed Table Information
Database:               retail
OwnerType:              USER
Owner:                  root
CreateTime:             Sun Oct 04 23:21:28 UTC 2026
LastAccessTime:         UNKNOWN
Retention:              0
Location:               hdfs://localhost:9000/user/hive/warehouse/retail.db/sales
Table Type:             MANAGED_TABLE
Table Parameters:
    COLUMN_STATS_ACCURATE   {\"BASIC_STATS\":\"true\"}
    bucketing_version       2
    numFiles                6
    numPartitions           2
    numRows                 9
    rawDataSize             72
    totalSize               6690
    transient_lastDdlTime   1791156088

# Storage Information
SerDe Library:          org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe
InputFormat:            org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat
OutputFormat:           org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat
Compressed:             No
Num Buckets:            3
Bucket Columns:         [store]
Sort Columns:           []
Storage Desc Params:
    serialization.format    1
South   10360.0 48
North   2520.0  39
{"input_tables":[{"tablename":"retail@sales","tabletype":"MANAGED_TABLE"}],"input_partitions":[{"partitionName":"retail@sales@quarter=Q2"}]}
Grocery Tea 500g    12  2520.0
Grocery Rice 5kg    4   1120.0
North   Rice 5kg    1120.0  1
North   Notebook    800.0   2
North   Notebook    600.0   3
South   Rice 5kg    2800.0  1
South   Tea 500g    2520.0  2
South   Tea 500g    1680.0  3
South   Rice 5kg    1680.0  3
South   Shampoo 200ml   980.0   5
South   Shampoo 200ml   700.0   6
$ awk '/^FAILED|^Error|Exception/' hive.log | head -5

The Python check, 10_hive_duckdb.py:

OUTPUT

  Experiment 10 -- Hive-style SQL over the star schema

    9 fact rows, total revenue 12,880

    region         revenue    profit   margin
    South           10,360     2,760   26.64%
    North            2,520       765   30.36%
         South = 10,360 is the SAME number Course 11's DAX
         CALCULATE measure produced. Two engines, two languages, one
         dataset -- if they ever disagree the suite fails, which is
         what makes the cross-check worth having

    partitioning by quarter -- what Hive actually does:
      partition       rows     revenue  scanned for Q2
      quarter=Q1          5       7,660               0
      quarter=Q2          4       5,220               4
      TOTAL              9                           4
         a partitioned table stores each quarter in its own HDFS
         DIRECTORY, so 'WHERE quarter = ''Q2''' reads 4 rows instead
         of 9 -- partition PRUNING, decided before a single byte is
         read. The partition column is a directory name, not a column
         in the data files, which is why it costs no storage

      the trap: partition on something with FEW distinct values.
      Partitioning by date_key here would make 4 directories for 9
      rows -- the small-files problem from experiment 4, created on
      purpose. Partition by quarter or month; BUCKET by customer_id

    bucketing (CLUSTERED BY store INTO 3 BUCKETS):
      bucket 0: ['Guntur', 'Hyderabad']
      bucket 1: (empty)
      bucket 2: ['Vijayawada']
         BUCKET 1 IS EMPTY, with three stores over three buckets.
         Hashing does not distribute small key sets evenly, and an
         empty bucket is still a file the job opens.
         Buckets are FILES inside a partition, assigned by a hash
         of the column. Two tables bucketed the same way on the same
         column can be joined bucket-to-bucket with no shuffle at all
         -- a sort-merge bucket join, and the reason bucketing exists

    managed against external tables:
                  data lives                DROP TABLE deletes
      MANAGED     /user/hive/warehouse      the DATA too
      EXTERNAL    wherever you point it     only the metadata
         use EXTERNAL for data you did not produce and cannot
         recreate. A DROP TABLE on a managed table over the company's
         only copy of a dataset is the classic Hive accident

    products above 1,000 revenue:
      category    product          units   revenue
      Grocery     Rice 5kg            20     5,600
      Grocery     Tea 500g            20     4,200
      Personal    Shampoo 200ml       12     1,680
      Stationery  Notebook            35     1,400
         HAVING filters GROUPS, WHERE filters ROWS -- and in Hive
         that distinction is a job-plan difference, not a syntax
         nicety: a WHERE on a partition column prunes directories
         before the job starts, a HAVING cannot

    what Hive is NOT:
      expectation                   reality
      row-level UPDATE/DELETE       only with ACID tables + ORC + buckets
      sub-second queries            seconds to minutes -- it plans a JOB
      indexes                       removed in Hive 3; use partitions and ORC/Parquet
      a running server holding data metadata only; data is files in HDFS
      enforced constraints          declarative only; NOT enforced
         Hive is a COMPILER: HiveQL in, a MapReduce/Tez/Spark job
         out. Everything surprising about it follows from that one
         sentence, and it is the right answer to 'compare Hive with an
         RDBMS'

PARTITION PRUNING

Partition Rows Revenue Scanned for Q2
quarter=Q1 5 7,660 0
quarter=Q2 4 5,220 4
total 9 12,880 4 of 9

A partition is an HDFS directory, so WHERE quarter='Q2' reads one directory — decided before a byte is read; Hive's EXPLAIN DEPENDENCY above lists the one partition it will read. The partition column is a directory name, so it costs no storage.

The trap: partitioning by date_key here would make 4 directories for 9 rows — the small-files problem, created on purpose.

BUCKETING, AND AN HONEST RESULT

CLUSTERED BY (store) INTO 3 BUCKETS over three stores:

bucket 0: ['Guntur', 'Hyderabad']
bucket 1: (empty)
bucket 2: ['Vijayawada']

Bucket 1 is empty. Hashing does not distribute small key sets evenly, and an empty bucket is still a file the job opens. Reporting that is worth more than pretending the hash was balanced.

Managed against external. DROP TABLE on a MANAGED table deletes the data. Use EXTERNAL for anything you did not produce and cannot recreate — the classic Hive accident.

HAVING against WHERE. Four products clear ₹1,000 (Grocery total ₹9,800). WHERE filters rows, HAVING filters groups — and in Hive that is a job-plan difference: a WHERE on a partition column prunes directories before the job starts, a HAVING cannot.

WHAT HIVE IS NOT

Expectation Reality
row-level UPDATE/DELETE only with ACID + ORC + buckets
sub-second queries seconds to minutes — it plans a job
indexes removed in Hive 3
a server holding data metadata only
enforced constraints declarative, not enforced

Hive is a compiler. Everything surprising follows from that sentence.

RESULT

Hive, on the cluster, gives South ₹10,360 from 48 units and North ₹2,520 from 39 — the number Business Intelligence Tools' DAX produced — and its query for Q2 reads one partition. DuckDB agrees on every figure.

Experiment 11 — Importing from a relational database with Sqoop

1. Question

Import data from an RDBMS into Hadoop using Sqoop.

2. Aim

Import a MySQL table into HDFS in parallel, import a query and a table into Hive, import only the new rows, save the import as a job, and export a summary back; then model the split queries Sqoop runs.

3. Steps

On the cluster, 11_sqoop.sh:

  1. Store the password.
  2. List the databases and tables.
  3. Import a table, in parallel.
  4. Import a query.
  5. Import into Hive.
  6. Import only what is new.
  7. Export back to the database.

The Python check, 11_sqoop_equivalent.py:

  1. Build the source table.
  2. Find the split boundaries.
  3. Split by a skewed column.
  4. Write the target.
  5. Import incrementally.

SPLITTING BY THE PRIMARY KEY

step 1: SELECT MIN(order_id), MAX(order_id) -> 1, 90
step 2: four ranges, one per mapper
Mapper WHERE Rows
0 order_id >= 1 AND <= 22 22
1 order_id >= 23 AND <= 45 23
2 order_id >= 46 AND <= 67 22
3 order_id >= 68 AND <= 90 23

Four TCP connections to the database. Sqoop's parallelism is database parallelism — -m 20 against a production OLTP box is a denial of service you wrote yourself. The Python half has a real SQLite database at one end and a real Parquet file at the other. The cluster's run prints the first step: BoundingValsQuery: SELECT MIN(order_id), MAX(order_id) FROM orders.

4. Programme

On the cluster, 11_sqoop.sh:

# Experiment 11 -- import data from an RDBMS into Hadoop using Sqoop
#
# Run it: bash 11_sqoop.sh, with MySQL and a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 11_sqoop_equivalent.py, which does the same import from a real SQLite database
#
# The database: MySQL (MariaDB) on this machine, with retail.orders -- 90 rows,
# ten copies of the nine shared sales rows -- and retail.customers.
# [Corrected: the commands connected to dbhost, a name for your database
# server; it is localhost here. Every command below ran.]

# Step 1: Store the password
echo -n "student-pw" > .pw
hdfs dfs -put .pw /user/student/.pw && rm .pw
hdfs dfs -chmod 400 /user/student/.pw
# [Changed: these three lines are added -- the commands below read the file.]

# Step 2: List the databases and tables
sqoop list-databases --connect jdbc:mysql://localhost:3306 \
                     --username student --password-file /user/student/.pw 2>/dev/null
sqoop list-tables --connect jdbc:mysql://localhost:3306/retail \
                  --username student --password-file /user/student/.pw 2>/dev/null
# NEVER use --password on the command line: it lands in `ps` and in the
# shell history. --password-file, on HDFS, mode 400.

# Step 3: Import a table, in parallel
#   --split-by order_id   Sqoop runs SELECT MIN/MAX on THIS column
#   --num-mappers 4       4 range queries, 4 connections
#   the null flags        without them, SQL NULL becomes the literal string
#                         "null" and every downstream count is wrong
sqoop import \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --table orders \
  --split-by order_id \
  --num-mappers 4 \
  --target-dir /user/student/orders \
  --as-parquetfile \
  --compress --compression-codec snappy \
  --null-string '\\N' --null-non-string '\\N' \
  2>&1 | grep -E "BoundingValsQuery|Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hdfs dfs -ls /user/student/orders | awk 'NR>1{print $NF}' | grep -v "^/user/student/orders/\.metadata"
# [Corrected: the comments sat after the backslashes, as `--split-by order_id \  # ...`.
# A backslash then escapes the space, not the newline, so the command ended at the
# comment and each following option ran as a command of its own -- `--num-mappers:
# command not found`. A comment cannot sit on a continued line.]

# Step 4: Import a query
#   $CONDITIONS is MANDATORY and not optional decoration: Sqoop substitutes
#   each mapper's range predicate there. Omit it and the import fails.
sqoop import \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --query 'SELECT o.*, c.region AS cust_region FROM orders o JOIN customers c
           ON o.cust_id = c.id WHERE $CONDITIONS' \
  --split-by o.order_id \
  --target-dir /user/student/orders_enriched \
  2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hdfs dfs -cat /user/student/orders_enriched/part-m-00000 | head -2
# [Corrected: the query was `SELECT o.*, c.region`. orders has a region column
# of its own, so the result had two columns named region, and Sqoop, which makes
# one Java field per column, stopped: "Import failed: Duplicate Column identifier
# specified: 'region'". Every column of a --query import needs its own name.]

# Step 5: Import into Hive
export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:$HIVE_HOME/lib/hive-common-3.1.3.jar
# [Corrected: Sqoop reads Hive's settings with Hive's own HiveConf class, and
# without that jar on HADOOP_CLASSPATH it fails with "ClassNotFoundException:
# org.apache.hadoop.hive.conf.HiveConf". Only that jar: with all of Hive's lib/,
# as is often advised, Sqoop runs Hive inside its own JVM, under a security
# manager it installs, and Hive's Derby metastore is refused -- "access denied
# org.apache.derby.security.SystemPermission( "engine", "usederbyinternals" )",
# ten retries, then failure. With just HiveConf it runs the hive command instead.]
sqoop import \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --table customers -m 1 \
  --hive-import --hive-database retail --hive-table customers \
  --create-hive-table 2>&1 | grep -E "Hive import complete|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hive -e "SELECT region, COUNT(*) FROM retail.customers GROUP BY region" 2>/dev/null | grep -v "^WARN"
# [Corrected: this was `sqoop import ... --hive-import`, with three dots where
# the connection and the table go; and a table with no primary key needs -m 1.]

# Step 6: Import only what is new
# ten new orders arrive in the database
mysql -h 127.0.0.1 -u student -pstudent-pw retail -e \
  "INSERT INTO orders SELECT order_id + 90, cust_id, store, region, product, category, qty, revenue, NOW()
   FROM orders WHERE order_id <= 10"
sqoop import \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --table orders --target-dir /user/student/orders_inc -m 1 \
  --incremental append --check-column order_id --last-value 90 \
  2>&1 | grep -E "Retrieved [0-9]+ records|--last-value|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
# lastmodified needs a timestamp column, and --merge-key so an updated row
# REPLACES its old copy rather than being added beside it:
#   sqoop import ... --incremental lastmodified --check-column updated_at \
#                    --last-value '2026-08-01 00:00:00' --merge-key order_id

# a saved job REMEMBERS --last-value for you
sqoop job --create orders_inc -- import \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --table orders --target-dir /user/student/orders_job -m 1 \
  --incremental append --check-column order_id --last-value 0 2>/dev/null
sqoop job --exec orders_inc 2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
sqoop job --show orders_inc 2>/dev/null | grep -E "incremental.last.value|incremental.col"
# [Corrected: these were `sqoop job --create orders_inc -- import ...`, with the
# connection and table left out as three dots.]
# [Noted: Sqoop 1.4.7 does not ship the org.json jar its saved jobs need. Without
# it, --create fails with "NoClassDefFoundError: org/json/JSONObject" yet keeps
# the job's name, with none of its options, and --exec then says "--table or
# --query is required for import". setup_hadoop.sh adds the jar.]

# Step 7: Export back to the database
#   an export is NOT transactional across mappers. If mapper 3 fails, the
#   rows mappers 1 and 2 wrote are already committed. Use --staging-table
#   when that matters.
mysql -h 127.0.0.1 -u student -pstudent-pw retail -B -N -e \
  "SELECT category, COUNT(*), SUM(qty), SUM(revenue) FROM orders GROUP BY category" \
  | tr '\t' ',' > by_category.csv
hdfs dfs -mkdir -p /user/student/out/by_category
hdfs dfs -put by_category.csv /user/student/out/by_category/
sqoop export \
  --connect jdbc:mysql://localhost:3306/retail \
  --username student --password-file /user/student/.pw \
  --table order_summary \
  --export-dir /user/student/out/by_category \
  --update-mode allowinsert --update-key category \
  --batch 2>&1 | grep -E "Exported [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
mysql -h 127.0.0.1 -u student -pstudent-pw retail -e "SELECT * FROM order_summary ORDER BY category"
# [Changed: the summary to export is made here from the database itself; the
# script assumed it was already on HDFS.]

# --- the three that bite ----------------------------------------------------
# 1. no primary key and no --split-by  -> Sqoop refuses; use -m 1
# 2. --split-by on a skewed column     -> one mapper does most of the work
# 3. neither mode notices a DELETE     -> re-import fully, periodically

The Python check, 11_sqoop_equivalent.py:

"""Experiment 11 -- import data from an RDBMS into Hadoop using Sqoop.

`11_sqoop.sh` carries the real commands, and runs with Sqoop against MySQL on a
cluster (the lab page shows it). What runs here is the same import in Python:
a REAL relational database (SQLite), a REAL split-by query, REAL parallel range
reads, and a REAL Parquet file at the other end.

Sqoop's entire trick is one line of SQL you never see:
    SELECT MIN(id), MAX(id) FROM table
and then one range query per mapper. Everything students find confusing about
Sqoop -- why --split-by matters, why a text primary key breaks it, why 4
mappers can produce wildly unequal files -- follows from that.
"""
import os
import sqlite3
import tempfile

import pyarrow as pa
import pyarrow.parquet as pq

import fixtures as f

FIELDS = ["order_id", "store", "region", "product", "category", "qty", "revenue"]


def build_rdbms(path, rows):
    con = sqlite3.connect(path)
    con.execute("""CREATE TABLE orders (
        order_id INTEGER PRIMARY KEY, store TEXT, region TEXT,
        product TEXT, category TEXT, qty INTEGER, revenue REAL)""")
    con.executemany("INSERT INTO orders VALUES (?,?,?,?,?,?,?)", rows)
    con.commit()
    return con


def boundaries(con, table, col, mappers):
    """Exactly what Sqoop runs before it launches a single mapper."""
    lo, hi = con.execute(f"SELECT MIN({col}), MAX({col}) FROM {table}").fetchone()
    step = (hi - lo + 1) / mappers
    out = []
    for m in range(mappers):
        a = lo + int(m * step)
        b = lo + int((m + 1) * step) - 1 if m < mappers - 1 else hi
        out.append((a, b))
    return lo, hi, out


def main():
    print("  Experiment 11 -- Sqoop import, with a real database at one end")

    tmp = tempfile.mkdtemp(prefix="bigdata11_")
    db = os.path.join(tmp, "retail.db")

    # Step 1: Build the source table
    # 90 orders, built from the nine shared rows so the totals stay checkable
    base = f.SALES_DF
    rows = []
    for i in range(90):
        r = base.iloc[i % 9]
        rows.append((i + 1, r["store"], r["region"], r["product"],
                     r["category"], int(r["qty"]), float(r["revenue"])))
    con = build_rdbms(db, rows)
    n = con.execute("SELECT COUNT(*) FROM orders").fetchone()[0]
    total = con.execute("SELECT SUM(revenue) FROM orders").fetchone()[0]
    print(f"\n    source: SQLite table 'orders', {n} rows, "
          f"revenue {total:,.0f}")
    assert n == 90 and abs(total - f.total_revenue() * 10) < 1e-6
    print(f"""         the 90 rows are ten copies of Course 11's nine, so the
         source total is exactly 10 x {f.total_revenue():,.0f}. If the imported
         Parquet does not carry that number, the import lost data --
         and that is the only import test that matters""")

    # Step 2: Find the split boundaries
    lo, hi, ranges = boundaries(con, "orders", "order_id", 4)
    print(f"\n    step 1: SELECT MIN(order_id), MAX(order_id) -> {lo}, {hi}")
    print(f"    step 2: split into 4 ranges, one per mapper")
    print(f"      {'mapper':<8}{'WHERE clause':<44}{'rows':>6}")
    imported = []
    for m, (a, b) in enumerate(ranges):
        where = f"order_id >= {a} AND order_id <= {b}"
        part = con.execute(
            f"SELECT {', '.join(FIELDS)} FROM orders WHERE {where}").fetchall()
        imported.extend(part)
        print(f"      {m:<8}{where:<44}{len(part):>6}")
    counts = [len(con.execute(
        f"SELECT 1 FROM orders WHERE order_id >= {a} AND order_id <= {b}"
    ).fetchall()) for a, b in ranges]
    assert sum(counts) == n
    assert max(counts) - min(counts) <= 1, "an integer key splits near-evenly"
    print(f"""         four mappers, {min(counts)} or {max(counts)} rows each ({n} does not
         divide by 4), four TCP connections to
         the database. Sqoop's parallelism is DATABASE parallelism --
         raise -m to 20 on a production OLTP box and you have written
         a denial of service against your own company""")

    # Step 3: Split by a skewed column
    print("\n    now split by a column that is NOT uniform -- 'qty':")
    lo2, hi2, ranges2 = boundaries(con, "orders", "qty", 4)
    print(f"      MIN(qty), MAX(qty) = {lo2}, {hi2}")
    print(f"      {'mapper':<8}{'range':<20}{'rows':>6}")
    skew = []
    for m, (a, b) in enumerate(ranges2):
        c = con.execute(
            f"SELECT COUNT(*) FROM orders WHERE qty >= {a} AND qty <= {b}"
        ).fetchone()[0]
        skew.append(c)
        print(f"      {m:<8}{f'{a}..{b}':<20}{c:>6}")
    assert sum(skew) == n
    assert max(skew) > 3 * min(skew) if min(skew) else True
    print(f"""         {max(skew)} rows for one mapper and {min(skew)} for another.
         Sqoop assumes the split column is UNIFORMLY DISTRIBUTED
         between its min and max, and qty is not. The job's wall
         clock is the slowest mapper, so a bad --split-by wastes
         three quarters of your parallelism.
         Split on the PRIMARY KEY unless you have measured otherwise""")

    print("\n    --split-by on a TEXT column:")
    print("      Sqoop needs an ORDERED, NUMERIC column to compute ranges.")
    print("      On text it must either refuse, or use")
    print("      -Dorg.apache.sqoop.splitter.allow_text_splitter=true")
    print("      which splits on string ordering and skews horribly.")
    print("""         a table with a UUID or composite primary key has no
         natural split column, and the honest answer is -m 1 --
         one mapper, no parallelism, correct results""")

    # Step 4: Write the target
    table = pa.Table.from_pylist(
        [dict(zip(FIELDS, r)) for r in sorted(imported)])
    out = os.path.join(tmp, "orders.parquet")
    pq.write_table(table, out, compression="snappy")
    back = pq.read_table(out)
    imported_total = sum(back.column("revenue").to_pylist())
    print(f"\n    landed: {out.split(os.sep)[-1]}, {back.num_rows} rows, "
          f"revenue {imported_total:,.0f}")
    assert back.num_rows == n
    assert abs(imported_total - total) < 1e-6, "the import must not lose money"
    print("""         row count AND the sum of a money column, both checked.
         Counting rows alone would not catch a truncated numeric
         type, which is the classic Sqoop bug: an Oracle NUMBER(38)
         silently becoming a Java double""")

    # Step 5: Import incrementally
    print("\n    incremental import, the two modes:")
    print(f"      {'mode':<15}{'--check-column':<18}{'catches'}")
    print(f"      {'append':<15}{'an increasing id':<18}"
          f"{'new rows only'}")
    print(f"      {'lastmodified':<15}{'a timestamp':<18}"
          f"{'new AND updated rows'}")
    last = con.execute("SELECT MAX(order_id) FROM orders").fetchone()[0]
    con.executemany("INSERT INTO orders VALUES (?,?,?,?,?,?,?)",
                    [(91, "Guntur", "South", "Tea 500g", "Grocery", 3, 630.0)])
    con.commit()
    new = con.execute(
        f"SELECT COUNT(*) FROM orders WHERE order_id > {last}").fetchone()[0]
    assert new == 1
    print(f"\n      --last-value {last} now selects {new} row")
    print("""         NEITHER MODE CATCHES A DELETE. Sqoop has no way to see a
         row that is gone, so an incrementally imported table drifts
         away from its source over time. The fix is a periodic full
         re-import, and knowing that is the difference between having
         used Sqoop and having read about it""")

    con.close()
    os.remove(out)
    os.remove(db)
    os.rmdir(tmp)


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 11_sqoop.sh:

OUTPUT

$ hdfs dfs -mkdir -p /user/student /user/hive/warehouse /tmp/hive
$ hdfs dfs -chmod 733 /tmp/hive && hdfs dfs -chmod g+w /user/hive/warehouse
$ schematool -dbType derby -initSchema 2>&1 | grep -E 'schemaTool completed'
schemaTool completed
$ hive -e 'CREATE DATABASE retail' 2>/dev/null | grep -v '^WARN'
$ echo -n "student-pw" > .pw
$ hdfs dfs -put .pw /user/student/.pw && rm .pw
$ hdfs dfs -chmod 400 /user/student/.pw
$ sqoop list-databases --connect jdbc:mysql://localhost:3306 \
                       --username student --password-file /user/student/.pw 2>/dev/null
information_schema
test
retail
$ sqoop list-tables --connect jdbc:mysql://localhost:3306/retail \
                    --username student --password-file /user/student/.pw 2>/dev/null
customers
order_summary
orders
$ sqoop import \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --table orders \
    --split-by order_id \
    --num-mappers 4 \
    --target-dir /user/student/orders \
    --as-parquetfile \
    --compress --compression-codec snappy \
    --null-string '\\N' --null-non-string '\\N' \
    2>&1 | grep -E "BoundingValsQuery|Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO db.DataDrivenDBInputFormat: BoundingValsQuery: SELECT MIN(`order_id`), MAX(`order_id`) FROM `orders`
INFO mapreduce.ImportJobBase: Retrieved 90 records.
$ hdfs dfs -ls /user/student/orders | awk 'NR>1{print $NF}' | grep -v "^/user/student/orders/\.metadata"
/user/student/orders/.signals
/user/student/orders/01ea0149-6607-4381-9487-10a9b64e080e.parquet
/user/student/orders/037af4d2-e6a3-43c3-98a4-c055218b97ec.parquet
/user/student/orders/24ad948c-16db-4ef2-bb18-3ce84fbfd6ad.parquet
/user/student/orders/df60c83e-1143-4106-80b7-ab7fc1e3a2fd.parquet
$ sqoop import \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --query 'SELECT o.*, c.region AS cust_region FROM orders o JOIN customers c
             ON o.cust_id = c.id WHERE $CONDITIONS' \
    --split-by o.order_id \
    --target-dir /user/student/orders_enriched \
    2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 90 records.
$ hdfs dfs -cat /user/student/orders_enriched/part-m-00000 | head -2
1,1,Vijayawada,South,Rice 5kg,Grocery,10,2800.0,2026-08-01 09:00:00.0,South
2,2,Vijayawada,South,Shampoo 200ml,Personal,5,700.0,2026-08-01 09:00:00.0,South
$ export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:$HIVE_HOME/lib/hive-common-3.1.3.jar
$ sqoop import \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --table customers -m 1 \
    --hive-import --hive-database retail --hive-table customers \
    --create-hive-table 2>&1 | grep -E "Hive import complete|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO hive.HiveImport: Hive import complete.
$ hive -e "SELECT region, COUNT(*) FROM retail.customers GROUP BY region" 2>/dev/null | grep -v "^WARN"
North   2
South   4
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -e \
    "INSERT INTO orders SELECT order_id + 90, cust_id, store, region, product, category, qty, revenue, NOW()
     FROM orders WHERE order_id <= 10"
$ sqoop import \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --table orders --target-dir /user/student/orders_inc -m 1 \
    --incremental append --check-column order_id --last-value 90 \
    2>&1 | grep -E "Retrieved [0-9]+ records|--last-value|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 10 records.
INFO tool.ImportTool:   --last-value 100
$ sqoop job --create orders_inc -- import \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --table orders --target-dir /user/student/orders_job -m 1 \
    --incremental append --check-column order_id --last-value 0 2>/dev/null
$ sqoop job --exec orders_inc 2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 100 records.
$ sqoop job --show orders_inc 2>/dev/null | grep -E "incremental.last.value|incremental.col"
incremental.last.value = 100
incremental.col = order_id
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -B -N -e \
    "SELECT category, COUNT(*), SUM(qty), SUM(revenue) FROM orders GROUP BY category" \
    | tr '\t' ',' > by_category.csv
$ hdfs dfs -mkdir -p /user/student/out/by_category
$ hdfs dfs -put by_category.csv /user/student/out/by_category/
$ sqoop export \
    --connect jdbc:mysql://localhost:3306/retail \
    --username student --password-file /user/student/.pw \
    --table order_summary \
    --export-dir /user/student/out/by_category \
    --update-mode allowinsert --update-key category \
    --batch 2>&1 | grep -E "Exported [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ExportJobBase: Exported 3 records.
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -e "SELECT * FROM order_summary ORDER BY category"
category    orders  units   revenue
Grocery 56  450 110600
Personal    22  132 18480
Stationery  22  385 15400

The Python check, 11_sqoop_equivalent.py:

OUTPUT

  Experiment 11 -- Sqoop import, with a real database at one end

    source: SQLite table 'orders', 90 rows, revenue 128,800
         the 90 rows are ten copies of Course 11's nine, so the
         source total is exactly 10 x 12,880. If the imported
         Parquet does not carry that number, the import lost data --
         and that is the only import test that matters

    step 1: SELECT MIN(order_id), MAX(order_id) -> 1, 90
    step 2: split into 4 ranges, one per mapper
      mapper  WHERE clause                                  rows
      0       order_id >= 1 AND order_id <= 22                22
      1       order_id >= 23 AND order_id <= 45               23
      2       order_id >= 46 AND order_id <= 67               22
      3       order_id >= 68 AND order_id <= 90               23
         four mappers, 22 or 23 rows each (90 does not
         divide by 4), four TCP connections to
         the database. Sqoop's parallelism is DATABASE parallelism --
         raise -m to 20 on a production OLTP box and you have written
         a denial of service against your own company

    now split by a column that is NOT uniform -- 'qty':
      MIN(qty), MAX(qty) = 4, 20
      mapper  range                 rows
      0       4..7                    40
      1       8..11                   20
      2       12..15                  20
      3       16..20                  10
         40 rows for one mapper and 10 for another.
         Sqoop assumes the split column is UNIFORMLY DISTRIBUTED
         between its min and max, and qty is not. The job's wall
         clock is the slowest mapper, so a bad --split-by wastes
         three quarters of your parallelism.
         Split on the PRIMARY KEY unless you have measured otherwise

    --split-by on a TEXT column:
      Sqoop needs an ORDERED, NUMERIC column to compute ranges.
      On text it must either refuse, or use
      -Dorg.apache.sqoop.splitter.allow_text_splitter=true
      which splits on string ordering and skews horribly.
         a table with a UUID or composite primary key has no
         natural split column, and the honest answer is -m 1 --
         one mapper, no parallelism, correct results

    landed: orders.parquet, 90 rows, revenue 128,800
         row count AND the sum of a money column, both checked.
         Counting rows alone would not catch a truncated numeric
         type, which is the classic Sqoop bug: an Oracle NUMBER(38)
         silently becoming a Java double

    incremental import, the two modes:
      mode           --check-column    catches
      append         an increasing id  new rows only
      lastmodified   a timestamp       new AND updated rows

      --last-value 90 now selects 1 row
         NEITHER MODE CATCHES A DELETE. Sqoop has no way to see a
         row that is gone, so an incrementally imported table drifts
         away from its source over time. The fix is a periodic full
         re-import, and knowing that is the difference between having
         used Sqoop and having read about it

The database is MariaDB, made by _drive_11_sqoop.py: retail.orders with 90 rows — ten copies of the nine sales rows, ₹128,800, exactly ten times Business Intelligence Tools' ₹12,880 — and retail.customers, with no primary key. Running the script found three faults, noted in it: comments after the line-continuing backslashes, which ended each command early; the query's two columns named region; and Hive's jars on the classpath, needed for the Hive import — but only its HiveConf, or the import fails inside Sqoop's own security manager. And Sqoop 1.4.7 does not ship the org.json jar its saved jobs need; setup_hadoop.sh adds it.

SPLITTING BY A SKEWED COLUMN

Mapper qty range Rows
0 4..7 40
1 8..11 20
2 12..15 20
3 16..20 10

Forty against ten. Sqoop assumes the split column is uniformly distributed between min and max. The job's wall clock is the slowest mapper, so a bad --split-by wastes three quarters of your parallelism.

The import is verified two ways. 90 rows and ₹128,800 both check out. Counting rows alone would not catch a truncated numeric type — an Oracle NUMBER(38) silently becoming a Java double is the classic Sqoop corruption bug.

NEITHER INCREMENTAL MODE CATCHES A DELETE

--last-value 90 selects the new rows — 10 on the cluster, 1 in the model. But Sqoop has no way to see a row that is gone, so an incrementally imported table drifts from its source. The fix is a periodic full re-import — and knowing that is the difference between having used Sqoop and having read about it.

RESULT

From MariaDB: 90 orders imported by four mappers, 90 rows of a join, the customers table into Hive — South 4, North 2 — 10 new rows incrementally, 100 by a saved job that then remembers 100, and 3 summary rows exported back. In the model, the four range queries split 22, 23, 22, 23, and a skewed column splits 40 against 10.

Experiment 12 — Capturing log data with Flume

1. Question

Capture and store log or streaming data using Flume.

2. Aim

Run a Flume agent that tails a web server's access log, adds headers with interceptors, routes server errors to an alert sink and the rest to HDFS; then model its channel and back-pressure.

3. Steps

On the cluster, 12_flume.conf:

  1. Name the agent's parts.
  2. Tail the log files.
  3. Add headers with interceptors.
  4. Route on a header.
  5. Set up the channels.
  6. Set up the sinks.
  7. Wire it together.

The Python check, 12_flume_equivalent.py:

  1. Intercept an event.
  2. Run a healthy agent.
  3. Slow the sink.
  4. Fix it.
  5. Compare the channel types.
  6. Read what the sink writes.
  7. Set the rollover.

THE INTERCEPTOR

headers {'host': '10.0.0.1', 'status': '200'}
body    (unchanged, 78 chars)

An interceptor adds headers and leaves the body alone. Headers are what a multiplexing selector routes on, so "send 500s to the alert sink" is a header rule, not code.

4. Programme

On the cluster, 12_flume.conf:

# Experiment 12 -- capture and store log/streaming data using Flume
#
# Run it: the flume-ng command below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 12_flume_equivalent.py, which runs the source/channel/sink semantics
#
# run with:
#   flume-ng agent --conf $FLUME_HOME/conf --conf-file 12_flume.conf \
#                  --name a1 -Dflume.root.logger=INFO,console

# Step 1: Name the agent's parts
a1.sources  = tailsrc
a1.channels = memch filech
a1.sinks    = hdfssink alertsink

# Step 2: Tail the log files
a1.sources.tailsrc.type              = TAILDIR
a1.sources.tailsrc.filegroups        = f1
a1.sources.tailsrc.filegroups.f1     = /tmp/flume-lab/logs/access.log.*
a1.sources.tailsrc.positionFile      = /tmp/flume-lab/taildir_position.json
#   TAILDIR, not `exec tail -F`: exec sources lose everything on restart
#   because there is no position file. This is the single most common
#   Flume data-loss bug.

# Step 3: Add headers with interceptors
a1.sources.tailsrc.interceptors                 = ts host regex
a1.sources.tailsrc.interceptors.ts.type         = timestamp
a1.sources.tailsrc.interceptors.host.type       = host
a1.sources.tailsrc.interceptors.host.useIP      = true
a1.sources.tailsrc.interceptors.regex.type      = regex_extractor
a1.sources.tailsrc.interceptors.regex.regex     = ^\\S+ \\S+ \\S+ \\[[^\\]]+\\] "\\S+ (\\S+)[^"]*" (\\d{3})
a1.sources.tailsrc.interceptors.regex.serializers            = s1 s2
a1.sources.tailsrc.interceptors.regex.serializers.s1.name    = path
a1.sources.tailsrc.interceptors.regex.serializers.s2.name    = status
# [Corrected: the regex was written with single backslashes, ^\S+ \S+ ... (\d{3}).
# Flume reads this file as Java properties, where a backslash escapes the next
# character and is dropped -- \S became S and \d became d -- so the regex matched
# no line, no event got a status header, and every one went to the default
# channel. In a properties file every regex backslash is doubled.]

# Step 4: Route on a header
a1.sources.tailsrc.selector.type           = multiplexing
a1.sources.tailsrc.selector.header         = status
a1.sources.tailsrc.selector.mapping.500    = filech
a1.sources.tailsrc.selector.default        = memch

# Step 5: Set up the channels
a1.channels.memch.type                  = memory
a1.channels.memch.capacity              = 10000
a1.channels.memch.transactionCapacity   = 1000
#   capacity is EVENTS, not bytes. A full channel blocks the source --
#   back-pressure, which is correct behaviour and looks like a hang.

a1.channels.filech.type                 = file
a1.channels.filech.checkpointDir        = /tmp/flume-lab/checkpoint
a1.channels.filech.dataDirs             = /tmp/flume-lab/data
#   file channel: survives a crash, roughly 10x slower. "We must not lose
#   events" and `type = memory` cannot both be true.

# Step 6: Set up the sinks
a1.sinks.hdfssink.type              = hdfs
a1.sinks.hdfssink.hdfs.path         = hdfs://localhost:9000/logs/dt=%Y-%m-%d/hr=%H
a1.sinks.hdfssink.hdfs.filePrefix   = access
a1.sinks.hdfssink.hdfs.fileType     = DataStream
a1.sinks.hdfssink.hdfs.writeFormat  = Text
a1.sinks.hdfssink.hdfs.batchSize    = 1000
# 10 min, not the 30 s default; one HDFS block; 0 = never roll on count
a1.sinks.hdfssink.hdfs.rollInterval = 600
a1.sinks.hdfssink.hdfs.rollSize     = 134217728
a1.sinks.hdfssink.hdfs.rollCount    = 0
#   THE DEFAULTS (30 s / 1 KB / 10 events) MANUFACTURE THE SMALL-FILES
#   PROBLEM. Left alone they produce a file every few seconds, each a few
#   hundred bytes, and the NameNode pays for every one. Always override.
# [Corrected: the three comments sat after the values, as `rollInterval = 600
# # 10 min`. A properties file has no end-of-line comments: the value was the
# whole rest of the line, the agent refused it -- "NumberFormatException: For
# input string: "600         # 10 min, not the 30 s default"" -- and dropped the
# HDFS sink, so nothing reached HDFS. A comment needs a line of its own.]

a1.sinks.alertsink.type             = logger
# [Corrected: the paths were a server's -- /var/log/apache2, /var/lib/flume, and
# a NameNode called nn. Here the agent tails a copy of the lab's access log in
# /tmp/flume-lab, keeps its state there, and writes to the NameNode on localhost.]

# Step 7: Wire it together
a1.sources.tailsrc.channels  = memch filech
a1.sinks.hdfssink.channel    = memch
a1.sinks.alertsink.channel   = filech

The Python check, 12_flume_equivalent.py:

"""Experiment 12 -- capture and store log/streaming data using Flume.

`12_flume.conf` carries the real agent configuration, and runs in Flume itself
(the lab page shows it). What runs here is the AGENT'S SEMANTICS: a source, a
channel with a bounded capacity, a sink with a batch size, and what actually
happens when the sink is slower than the source -- which is the only Flume
question worth asking.
"""
from collections import deque

import fixtures as f


class Channel:
    """A bounded buffer. Flume's memory channel, minus the threads.

    The capacity is the whole story: a full channel makes the SOURCE block,
    which is back-pressure, which is the correct behaviour and the one that
    surprises people.
    """

    def __init__(self, capacity):
        self.capacity = capacity
        self.q = deque()
        self.rejected = 0
        self.high_water = 0

    def put(self, event):
        if len(self.q) >= self.capacity:
            self.rejected += 1
            return False
        self.q.append(event)
        self.high_water = max(self.high_water, len(self.q))
        return True

    def take(self, n):
        out = [self.q.popleft() for _ in range(min(n, len(self.q)))]
        return out


def interceptor(event):
    """Flume interceptors add HEADERS; they do not change the body."""
    ip = event.split(" ", 1)[0]
    status = event.rsplit(" ", 2)[-2]
    return {"headers": {"host": ip, "status": status}, "body": event}


def run_agent(events, capacity, batch, sink_every):
    """Drive one source -> channel -> sink agent, tick by tick."""
    chan = Channel(capacity)
    delivered, tick = [], 0
    src = list(events)
    while src or chan.q:
        if src:
            chan.put(interceptor(src.pop(0)))
        if tick % sink_every == 0:
            delivered.extend(chan.take(batch))
        tick += 1
    delivered.extend(chan.take(len(chan.q)))
    return chan, delivered, tick


def main():
    print("  Experiment 12 -- a Flume agent, semantics first")

    logs = f.access_logs(40)
    print(f"\n    source: {len(logs)} access-log lines")
    print(f"      {logs[0]}")

    # Step 1: Intercept an event
    ev = interceptor(logs[0])
    print(f"\n    after the interceptor:")
    print(f"      headers {ev['headers']}")
    print(f"      body    (unchanged, {len(ev['body'])} chars)")
    assert ev["headers"]["host"] == "10.0.0.1"
    assert ev["headers"]["status"] == "200"
    assert ev["body"] == logs[0]
    print("""         an interceptor adds HEADERS and leaves the body alone.
         Headers are what a multiplexing channel selector routes on,
         so 'send 500s to the alert sink and everything else to HDFS'
         is a header rule, not code""")

    # Step 2: Run a healthy agent
    chan, delivered, ticks = run_agent(logs, capacity=100, batch=10, sink_every=1)
    print(f"\n    capacity 100, batch 10, sink every tick:")
    print(f"      delivered {len(delivered)} of {len(logs)}, "
          f"rejected {chan.rejected}, peak channel depth {chan.high_water}")
    assert len(delivered) == len(logs) and chan.rejected == 0
    assert chan.high_water <= 10

    # Step 3: Slow the sink
    chan2, delivered2, _ = run_agent(logs, capacity=8, batch=4, sink_every=6)
    print(f"\n    capacity 8, batch 4, sink every 6th tick (a SLOW sink):")
    print(f"      delivered {len(delivered2)} of {len(logs)}, "
          f"rejected {chan2.rejected}, peak depth {chan2.high_water}")
    assert chan2.rejected > 0, "a slow sink must fill the channel"
    assert chan2.high_water == 8
    print(f"""         {chan2.rejected} events were REFUSED by the channel, because the
         sink could not drain it. In a real agent the source then
         BLOCKS rather than dropping -- back-pressure travels back up
         the pipe to the web server. 'Flume lost my events' almost
         always means 'the channel was full and the source gave up'""")

    # Step 4: Fix it
    print("\n    the fix, and its cost:")
    for cap in (8, 20, 100):
        c, d, _ = run_agent(logs, capacity=cap, batch=4, sink_every=6)
        print(f"      capacity {cap:>4}: rejected {c.rejected:>3}, "
              f"peak depth {c.high_water:>3}")
    print("""         a bigger channel absorbs a longer burst and buys nothing
         if the sink is permanently slower than the source. Buffers
         smooth BURSTS; they cannot fix a throughput deficit, and
         that sentence answers most Flume tuning questions""")

    # Step 5: Compare the channel types
    print("\n    channel types, and what you are choosing between:")
    print(f"      {'channel':<12}{'survives a crash?':<20}{'throughput'}")
    for name, durable, tput in (
            ("memory", "NO -- events lost", "highest"),
            ("file", "yes -- WAL on disk", "roughly 10x slower"),
            ("Kafka", "yes -- replicated", "high, but another cluster")):
        print(f"      {name:<12}{durable:<20}{tput}")
    print("""         a memory channel plus 'we must not lose events' is a
         contradiction, and it is the most common Flume misconfig.
         Choose the channel from the durability requirement, then
         size the cluster for whatever throughput that leaves""")

    # Step 6: Read what the sink writes
    from collections import Counter
    by_status = Counter(e["headers"]["status"] for e in delivered)
    by_host = Counter(e["headers"]["host"] for e in delivered)
    print(f"\n    what landed in HDFS, by header:")
    print(f"      status: {dict(sorted(by_status.items()))}")
    print(f"      host  : {dict(sorted(by_host.items()))}")
    assert sum(by_status.values()) == 40
    assert by_status["200"] == 24 and by_status["404"] == 8
    assert len(by_host) == 4 and all(v == 10 for v in by_host.values())
    print("""         24 successes, 8 not-founds and 8 server errors, evenly
         over four hosts. That breakdown is what experiment 17 reads
         back with Spark -- ingestion and analysis on the same bytes,
         which is the point of building the pipeline at all""")

    # Step 7: Set the rollover
    print("\n    the HDFS sink's rollover settings:")
    print(f"      {'setting':<24}{'default':<12}{'what it does'}")
    for k, v, w in (("hdfs.rollInterval", "30 sec", "close the file on a timer"),
                    ("hdfs.rollSize", "1024 bytes", "close it at a size"),
                    ("hdfs.rollCount", "10 events", "close it after N events")):
        print(f"      {k:<24}{v:<12}{w}")
    files = -(-len(logs) // 10)
    print(f"\n      at the DEFAULTS, {len(logs)} events produce ~{files} HDFS files")
    assert files == 4
    print("""         and every one of them is a few hundred bytes. Left alone,
         Flume's defaults manufacture the small-files problem from
         experiment 4 at a rate of two per minute. Set rollCount to 0
         and rollSize to a block, or run a compaction job -- this is
         the single most common Flume-in-production mistake""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 12_flume.conf:

OUTPUT

$ rm -rf /tmp/flume-lab && mkdir -p /tmp/flume-lab/logs
$ cp access.log /tmp/flume-lab/logs/access.log.1
$ wc -l < /tmp/flume-lab/logs/access.log.1
40
$ timeout 45 flume-ng agent --conf $FLUME_HOME/conf --conf-file 12_flume.conf --name a1 -Dflume.root.logger=INFO,console > agent.log 2>&1
$ grep -c 'LoggerSink: Event:' agent.log
8
$ grep 'LoggerSink: Event:' agent.log | grep -oE 'path=[^,]+|status=[0-9]+' | paste - - | sort | uniq -c
      8 path=/static/app.js status=500
$ hdfs dfs -ls -R /logs | awk '{print $NF}'
/logs/dt=2026-10-04
/logs/dt=2026-10-04/hr=23
/logs/dt=2026-10-04/hr=23/access.1791156576815
$ hdfs dfs -cat '/logs/*/*/*' | wc -l
32
$ hdfs dfs -cat '/logs/*/*/*' | awk '{print $9}' | sort | uniq -c
     24 200
      8 404
$ awk '/ (ERROR|FATAL) /' agent.log | sed -E 's/^[0-9T:,-]+ //' | sort -u | head -5

The Python check, 12_flume_equivalent.py:

OUTPUT

  Experiment 12 -- a Flume agent, semantics first

    source: 40 access-log lines
      10.0.0.1 - - [12/Aug/2025:09:00:00 +0530] "GET /index.html HTTP/1.1" 200 512

    after the interceptor:
      headers {'host': '10.0.0.1', 'status': '200'}
      body    (unchanged, 76 chars)
         an interceptor adds HEADERS and leaves the body alone.
         Headers are what a multiplexing channel selector routes on,
         so 'send 500s to the alert sink and everything else to HDFS'
         is a header rule, not code

    capacity 100, batch 10, sink every tick:
      delivered 40 of 40, rejected 0, peak channel depth 1

    capacity 8, batch 4, sink every 6th tick (a SLOW sink):
      delivered 32 of 40, rejected 8, peak depth 8
         8 events were REFUSED by the channel, because the
         sink could not drain it. In a real agent the source then
         BLOCKS rather than dropping -- back-pressure travels back up
         the pipe to the web server. 'Flume lost my events' almost
         always means 'the channel was full and the source gave up'

    the fix, and its cost:
      capacity    8: rejected   8, peak depth   8
      capacity   20: rejected   0, peak depth  16
      capacity  100: rejected   0, peak depth  16
         a bigger channel absorbs a longer burst and buys nothing
         if the sink is permanently slower than the source. Buffers
         smooth BURSTS; they cannot fix a throughput deficit, and
         that sentence answers most Flume tuning questions

    channel types, and what you are choosing between:
      channel     survives a crash?   throughput
      memory      NO -- events lost   highest
      file        yes -- WAL on disk  roughly 10x slower
      Kafka       yes -- replicated   high, but another cluster
         a memory channel plus 'we must not lose events' is a
         contradiction, and it is the most common Flume misconfig.
         Choose the channel from the durability requirement, then
         size the cluster for whatever throughput that leaves

    what landed in HDFS, by header:
      status: {'200': 24, '404': 8, '500': 8}
      host  : {'10.0.0.1': 10, '10.0.0.2': 10, '10.0.0.3': 10, '10.0.0.4': 10}
         24 successes, 8 not-founds and 8 server errors, evenly
         over four hosts. That breakdown is what experiment 17 reads
         back with Spark -- ingestion and analysis on the same bytes,
         which is the point of building the pipeline at all

    the HDFS sink's rollover settings:
      setting                 default     what it does
      hdfs.rollInterval       30 sec      close the file on a timer
      hdfs.rollSize           1024 bytes  close it at a size
      hdfs.rollCount          10 events   close it after N events

      at the DEFAULTS, 40 events produce ~4 HDFS files
         and every one of them is a few hundred bytes. Left alone,
         Flume's defaults manufacture the small-files problem from
         experiment 4 at a rate of two per minute. Set rollCount to 0
         and rollSize to a block, or run a compaction job -- this is
         the single most common Flume-in-production mistake

_drive_12_flume.py puts the lab's 40 access-log lines where the agent looks, runs it for 45 seconds, and counts what reached each sink. Running the configuration found two faults, both from its being a Java properties file:

THE CHANNEL, AND BACK-PRESSURE

Configuration Delivered Rejected Peak depth
capacity 100, batch 10, fast sink 40 0 ≤10
capacity 8, batch 4, slow sink 40 8 8
capacity 20, slow sink 40 0 16
capacity 100, slow sink 40 0 16

Eight events refused because the sink could not drain the channel. In a real agent the source then blocks — back-pressure travelling back to the web server. "Flume lost my events" almost always means "the channel was full and the source gave up". And a bigger channel absorbs a longer burst and buys nothing once the sink is permanently slower. Buffers smooth bursts; they cannot fix a throughput deficit.

What landed. {'200': 24, '404': 8, '500': 8}, evenly over four hosts (10 each) — the agent on the cluster split them the same way. Those same numbers appear in experiments 14 and 17 — three code paths, one set of figures.

THE DEFAULTS MANUFACTURE THE SMALL-FILES PROBLEM

rollInterval 30 s, rollSize 1024 bytes, rollCount 10 events → 40 events produce ~4 HDFS files, each a few hundred bytes. Left alone, Flume generates the Unit 2 small-files problem at two files a minute. With the configuration's settings the 32 events made one file.

RESULT

The agent read the 40 log lines, sent the 8 with status 500 to the logger sink and the other 32 — 24 of 200, 8 of 404 — to one HDFS file under dt=/hr=. Before two corrections it delivered nothing to HDFS and routed nothing.

Experiment 13 — Avro and Parquet

1. Question

Serialize and store datasets in Avro and Parquet formats.

2. Aim

Write the sales in Avro and in Parquet, read them back, evolve the Avro schema, read only some Parquet columns, and compare the sizes honestly.

3. Steps

The Python check, 13_avro_parquet.py:

  1. Write and read Avro.
  2. Evolve the schema.
  3. Write Parquet, and read only some columns.
  4. Compare the two formats.
  5. Test the claim at size.

THE REAL FORMATS

fastavro and pyarrow are real implementations, so the files written here are byte-for-byte readable by Hadoop, Hive and Spark.

Avro: self-describing. 9 records, 938 bytes, round-trip exact. The writer's schema is in the file header — full name in.ac.datascience.sales.Sale, namespace included.

4. Programme

The Python check, 13_avro_parquet.py:

"""Experiment 13 -- serialize and store datasets in Avro and Parquet.

THIS EXPERIMENT FULLY RUNS. fastavro and pyarrow are real implementations of
the real formats, so the files written here are byte-for-byte readable by
Hadoop, Hive and Spark. Nothing is simulated.

The point of the experiment is not "how do I call the library" -- it is the
difference between a ROW format and a COLUMN format, which decides everything
about how a big-data query performs.
"""
import io
import json
import os
import tempfile

import fastavro
import pyarrow as pa
import pyarrow.parquet as pq

import fixtures as f

AVRO_SCHEMA = {
    "type": "record",
    "name": "Sale",
    "namespace": "in.ac.datascience.sales",
    "fields": [
        {"name": "date_key", "type": "string"},
        {"name": "store", "type": "string"},
        {"name": "region", "type": "string"},
        {"name": "product", "type": "string"},
        {"name": "category", "type": "string"},
        {"name": "qty", "type": "long"},
        {"name": "revenue", "type": "double"},
        {"name": "profit", "type": "double"},
    ],
}

FIELDS = [fld["name"] for fld in AVRO_SCHEMA["fields"]]


def records():
    return [{k: (int(r[k]) if k == "qty" else r[k]) for k in FIELDS}
            for _, r in f.SALES_DF.iterrows()]


def main():
    print("  Experiment 13 -- Avro and Parquet, both really written")

    rows = records()
    tmp = tempfile.mkdtemp(prefix="bigdata13_")

    # Step 1: Write and read Avro
    avro_path = os.path.join(tmp, "sales.avro")
    with open(avro_path, "wb") as fh:
        fastavro.writer(fh, fastavro.parse_schema(AVRO_SCHEMA), rows)
    with open(avro_path, "rb") as fh:
        back = list(fastavro.reader(fh))
    assert back == rows, "Avro must round-trip exactly"
    avro_size = os.path.getsize(avro_path)
    # The notes quote this figure, so it is asserted. It is deterministic --
    # fixed rows, fixed schema -- but NOT independent of the schema's text:
    # the namespace is stored in the file header, so renaming it moves the
    # byte count. That is how this assertion earns its place; the figure had
    # already drifted once, silently, when the namespace changed.
    assert avro_size == 938, (
        f"Avro file is {avro_size} bytes, the notes say 938 -- "
        "update lab.md and unit-4.md, or find out what changed")
    print(f"\n    Avro   : {len(rows)} records, {avro_size} bytes, round-trip exact")

    # the schema travels INSIDE the file -- this is the property that matters
    with open(avro_path, "rb") as fh:
        embedded = fastavro.reader(fh).writer_schema
    assert embedded["name"] == "in.ac.datascience.sales.Sale", (
        "Avro stores the FULL name -- namespace + name -- not the short one")
    assert [fl["name"] for fl in embedded["fields"]] == FIELDS
    print(f"\n      the embedded schema's full name is {embedded['name']!r}")
    print("""         the WRITER'S SCHEMA is stored in the file header, so an
         Avro file is self-describing. A reader five years later needs
         no external metadata, which is exactly what a CSV cannot
         promise -- and why Avro is the ingestion format""")

    # Step 2: Evolve the schema
    evolved = json.loads(json.dumps(AVRO_SCHEMA))
    evolved["fields"].append(
        {"name": "channel", "type": ["null", "string"], "default": None})
    buf = io.BytesIO()
    with open(avro_path, "rb") as fh:
        old_bytes = fh.read()
    read_new = list(fastavro.reader(io.BytesIO(old_bytes),
                                    reader_schema=fastavro.parse_schema(evolved)))
    assert all(r["channel"] is None for r in read_new)
    assert len(read_new) == len(rows)
    print(f"\n    schema evolution: read {len(rows)} OLD records with a NEW schema")
    print(f"      the added field 'channel' comes back as "
          f"{read_new[0]['channel']!r} -- its DEFAULT")
    print("""         the old file was NOT rewritten. Avro resolves the writer's
         schema against the reader's, field by field, and fills in
         defaults for anything missing. A field added WITHOUT a
         default breaks exactly this, which is the one rule to
         remember about evolving an Avro schema""")

    # Step 3: Write Parquet, and read only some columns
    table = pa.Table.from_pylist(rows)
    pq_path = os.path.join(tmp, "sales.parquet")
    pq.write_table(table, pq_path, compression="snappy")
    pq_size = os.path.getsize(pq_path)
    back_pq = pq.read_table(pq_path).to_pylist()
    assert back_pq == rows, "Parquet must round-trip exactly"
    print(f"\n    Parquet: {len(rows)} records, {pq_size} bytes, round-trip exact")

    # column projection -- the whole reason Parquet exists
    one_col = pq.read_table(pq_path, columns=["revenue"])
    assert one_col.num_columns == 1 and one_col.num_rows == len(rows)
    meta = pq.ParquetFile(pq_path).metadata
    rg = meta.row_group(0)
    col_sizes = {rg.column(i).path_in_schema:
                 rg.column(i).total_compressed_size
                 for i in range(rg.num_columns)}
    print(f"\n    bytes stored PER COLUMN inside the Parquet file:")
    for name, sz in sorted(col_sizes.items(), key=lambda kv: -kv[1]):
        print(f"      {name:<12}{sz:>7}")
    total = sum(col_sizes.values())
    rev = col_sizes["revenue"]
    print(f"      {'TOTAL':<12}{total:>7}")
    print(f"\n    SELECT revenue reads {rev} of {total} column bytes "
          f"({100 * rev / total:.1f}%)")
    assert rev < total / 4
    print("""         THAT is column projection, and it is why Parquet wins on
         analytical queries: a SELECT of one column out of eight
         reads roughly one column's worth of bytes. A row format has
         to read every row in full and discard seven fields""")

    # predicate pushdown via row-group statistics
    stats = rg.column([i for i in range(rg.num_columns)
                       if rg.column(i).path_in_schema == "revenue"][0]).statistics
    print(f"\n    row-group statistics for 'revenue': "
          f"min {stats.min:,.0f}, max {stats.max:,.0f}")
    assert stats.min == 600.0 and stats.max == 2800.0
    print("""         a query for revenue > 5000 can SKIP THIS ENTIRE ROW GROUP
         without decoding a byte, because the max is 2,800. That is
         predicate pushdown, and on a partitioned Parquet dataset it
         is often a bigger win than the compression""")

    # Step 4: Compare the two formats
    csv_path = os.path.join(tmp, "sales.csv")
    f.SALES_DF[FIELDS].to_csv(csv_path, index=False)
    csv_size = os.path.getsize(csv_path)
    print(f"\n    {'format':<12}{'bytes':>8}  {'layout':<8}{'schema':<15}{'best for'}")
    for name, size, layout, schema, use in (
            ("CSV", csv_size, "row", "none", "interchange, and nothing else"),
            ("Avro", avro_size, "row", "in the file", "ingestion, streaming, evolution"),
            ("Parquet", pq_size, "COLUMN", "in the footer", "analytics, column projection"),
            ("SequenceFile", None, "row", "external", "legacy Hadoop key/value")):
        shown = f"{size:>8}" if size else f"{'--':>8}"
        print(f"    {name:<12}{shown}  {layout:<8}{schema:<15}{use}")
    print(f"""
         on NINE ROWS Parquet is LARGER than CSV ({pq_size} against
         {csv_size}) -- the footer, the schema and the per-column
         metadata are fixed overhead that nine rows cannot amortise.
         Report that honestly: Parquet's advantage is asymptotic, and
         quoting a compression ratio from a toy file is how people
         get caught out in a viva""")

    # Step 5: Test the claim at size
    # 108,000 records. TWO versions: one that repeats the nine rows exactly,
    # and one where every row differs -- because a columnar format's headline
    # ratio is mostly a statement about how repetitive the data is, and
    # quoting the repetitive number alone would be misleading.
    big = rows * 12000
    varied = [dict(r, qty=r["qty"] + i % 97,
                   revenue=r["revenue"] + (i % 8191) * 0.25,
                   store=f"{r['store']}-{i % 500}")
              for i, r in enumerate(big)]
    vsizes = {}
    for name, data in (("repetitive", big), ("varied", varied)):
        va = os.path.join(tmp, f"v_{name}.avro")
        vp = os.path.join(tmp, f"v_{name}.parquet")
        vc = os.path.join(tmp, f"v_{name}.csv")
        with open(va, "wb") as fh:
            fastavro.writer(fh, fastavro.parse_schema(AVRO_SCHEMA), data)
        pq.write_table(pa.Table.from_pylist(data), vp, compression="snappy")
        with open(vc, "w") as fh:
            fh.write(",".join(FIELDS) + "\n")
            for r in data:
                fh.write(",".join(str(r[k]) for k in FIELDS) + "\n")
        vsizes[name] = {"CSV": os.path.getsize(vc), "Avro": os.path.getsize(va),
                        "Parquet": os.path.getsize(vp)}
        for pth in (va, vp, vc):
            os.remove(pth)

    print(f"\n    the same schema at {len(big):,} records:")
    print(f"      {'data':<12}{'CSV':>12}{'Avro':>12}{'Parquet':>12}"
          f"{'CSV/Parquet':>13}")
    for name in ("repetitive", "varied"):
        z = vsizes[name]
        print(f"      {name:<12}{z['CSV']:>12,}{z['Avro']:>12,}"
              f"{z['Parquet']:>12,}{z['CSV'] / z['Parquet']:>12.1f}x")
    rep = vsizes["repetitive"]["CSV"] / vsizes["repetitive"]["Parquet"]
    var = vsizes["varied"]["CSV"] / vsizes["varied"]["Parquet"]
    assert rep > var * 5, "the repetitive figure must be visibly inflated"
    assert var > 1.5, "Parquet should still beat CSV on varied data"
    print(f"""         READ BOTH ROWS. The {rep:.0f}x on repetitive data is an
         ARTEFACT: 12,000 identical copies of nine rows dictionary-
         encode to almost nothing. Give every row a distinct store
         and revenue and the ratio falls to {var:.1f}x. Even that is
         optimistic -- date, region and category are still repetitive
         here -- and Parquet-against-CSV in production usually lands
         between 3x and 10x.
         A columnar format's headline compression number is mostly a
         statement about how REPETITIVE your data is, and a benchmark
         on duplicated rows says nothing at all""")

    print("\n    which format for which job:")
    print("      row-by-row WRITES, whole-record reads     -> Avro")
    print("      column aggregates over billions of rows   -> Parquet")
    print("      a landing zone that must survive schema")
    print("      changes for years                         -> Avro")
    print("      the table Hive and Spark actually query   -> Parquet")
    print("""         the standard architecture uses BOTH: Avro at the edge
         where records arrive one at a time and schemas drift, then a
         batch job converts to Parquet for the query layer. That
         answer is worth full marks on 'compare Avro and Parquet'""")

    for path in (avro_path, pq_path, csv_path):
        os.remove(path)
    os.rmdir(tmp)


if __name__ == "__main__":
    main()

5. Execution and Results

The Python check, 13_avro_parquet.py:

OUTPUT

  Experiment 13 -- Avro and Parquet, both really written

    Avro   : 9 records, 938 bytes, round-trip exact

      the embedded schema's full name is 'in.ac.datascience.sales.Sale'
         the WRITER'S SCHEMA is stored in the file header, so an
         Avro file is self-describing. A reader five years later needs
         no external metadata, which is exactly what a CSV cannot
         promise -- and why Avro is the ingestion format

    schema evolution: read 9 OLD records with a NEW schema
      the added field 'channel' comes back as None -- its DEFAULT
         the old file was NOT rewritten. Avro resolves the writer's
         schema against the reader's, field by field, and fills in
         defaults for anything missing. A field added WITHOUT a
         default breaks exactly this, which is the one rule to
         remember about evolving an Avro schema

    Parquet: 9 records, 2584 bytes, round-trip exact

    bytes stored PER COLUMN inside the Parquet file:
      profit          150
      qty             144
      revenue         143
      product         125
      category        111
      store           110
      date_key         83
      region           83
      TOTAL           949

    SELECT revenue reads 143 of 949 column bytes (15.1%)
         THAT is column projection, and it is why Parquet wins on
         analytical queries: a SELECT of one column out of eight
         reads roughly one column's worth of bytes. A row format has
         to read every row in full and discard seven fields

    row-group statistics for 'revenue': min 600, max 2,800
         a query for revenue > 5000 can SKIP THIS ENTIRE ROW GROUP
         without decoding a byte, because the max is 2,800. That is
         predicate pushdown, and on a partitioned Parquet dataset it
         is often a bigger win than the compression

    format         bytes  layout  schema         best for
    CSV              533  row     none           interchange, and nothing else
    Avro             938  row     in the file    ingestion, streaming, evolution
    Parquet         2584  COLUMN  in the footer  analytics, column projection
    SequenceFile      --  row     external       legacy Hadoop key/value

         on NINE ROWS Parquet is LARGER than CSV (2584 against
         533) -- the footer, the schema and the per-column
         metadata are fixed overhead that nine rows cannot amortise.
         Report that honestly: Parquet's advantage is asymptotic, and
         quoting a compression ratio from a toy file is how people
         get caught out in a viva

    the same schema at 108,000 records:
      data                 CSV        Avro     Parquet  CSV/Parquet
      repetitive     5,700,058   5,924,195      18,790       303.4x
      varied         6,269,519   6,380,511     522,264        12.0x
         READ BOTH ROWS. The 303x on repetitive data is an
         ARTEFACT: 12,000 identical copies of nine rows dictionary-
         encode to almost nothing. Give every row a distinct store
         and revenue and the ratio falls to 12.0x. Even that is
         optimistic -- date, region and category are still repetitive
         here -- and Parquet-against-CSV in production usually lands
         between 3x and 10x.
         A columnar format's headline compression number is mostly a
         statement about how REPETITIVE your data is, and a benchmark
         on duplicated rows says nothing at all

    which format for which job:
      row-by-row WRITES, whole-record reads     -> Avro
      column aggregates over billions of rows   -> Parquet
      a landing zone that must survive schema
      changes for years                         -> Avro
      the table Hive and Spark actually query   -> Parquet
         the standard architecture uses BOTH: Avro at the edge
         where records arrive one at a time and schemas drift, then a
         batch job converts to Parquet for the query layer. That
         answer is worth full marks on 'compare Avro and Parquet'

SCHEMA EVOLUTION, DEMONSTRATED

read 9 OLD records with a NEW nine-field schema
the added field 'channel' comes back as None -- its DEFAULT

The old file was not rewritten. Avro resolves writer's schema against reader's, field by field. A field added without a default breaks exactly this, and that is the one rule to remember.

PARQUET: COLUMN PROJECTION

Column Bytes
profit 150
qty 144
revenue 143
product 125
category 111
store 110
date_key 83
region 83
total 949

SELECT revenue reads 143 of 949 column bytes — 15.1%.

Predicate pushdown: row-group statistics for revenue are min 600, max 2,800, so a query for revenue > 5000 skips the entire row group without decoding a byte.

THE COMPRESSION CLAIM, TOLD HONESTLY

Data CSV Avro Parquet CSV/Parquet
9 rows 533 938 2,584 0.2×
108,000 rows, repetitive 5,700,058 5,924,190 18,790 303.4×
108,000 rows, varied 6,269,519 6,380,506 522,264 12.0×

On nine rows Parquet is 4.8× LARGER than CSV. The 303× is an artefact of 12,000 identical copies. Give every row a distinct store and revenue and it falls to 12.0× — and even that is optimistic. In production, 3× to 10×.

A columnar format's headline ratio is mostly a statement about how repetitive your data is, and a benchmark on duplicated rows says nothing at all.

RESULT

938 bytes of Avro and an exact round trip; old records read with a new schema get the added field's default; SELECT revenue reads 143 of 949 Parquet column bytes. On nine rows Parquet is 4.8 times larger than CSV, and the 303× ratio holds only for repetitive data — 12.0× once the rows vary.

Experiment 14 — An end-to-end ingestion workflow, batch and streaming

1. Question

Build an end-to-end ingestion workflow combining batch (Sqoop) and streaming (Flume).

2. Aim

Import the orders in a batch leg and the log events in a stream leg, both to Parquet, then join them in DuckDB — first wrongly, then at a common grain.

3. Steps

The Python check, 14_pipeline.py:

  1. Run the batch leg and the stream leg.
  2. Join them.
  3. Fix the join.
  4. Compare the two legs.
  5. Set lambda against kappa.
  6. Reconcile.

THE FAN TRAP

SQLite → Parquet (batch), log events → Parquet (streaming), then a real DuckDB query across both. The batch leg is experiment 11's import and the stream leg experiment 12's agent, each done here in Python so the two legs can be joined in one program; experiments 11 and 12 ran the real tools on the cluster.

events counted through the join: 90
events actually ingested       : 40

Each host appears in several orders, so every event is counted once per matching order. This is the same fan trap Business Intelligence Tools found in a Power BI model — not a SQL problem, a grain problem, appearing wherever two fact tables are joined directly.

4. Programme

The Python check, 14_pipeline.py:

"""Experiment 14 -- an end-to-end ingestion workflow combining batch (Sqoop)
and streaming (Flume).

This is the experiment that ties the course together, and the one where the
interesting problem is not any single tool but the JOIN BETWEEN THEM: batch
data arrives hourly and complete, streaming data arrives continuously and
incomplete, and a query that spans both has to decide what it means by "now".

Runs end to end: SQLite -> Parquet (the batch side), log events -> Parquet
(the streaming side), then a real DuckDB query across the two.
"""
import os
import sqlite3
import tempfile
from collections import Counter

import duckdb
import pyarrow as pa
import pyarrow.parquet as pq

import fixtures as f


def batch_leg(tmp):
    """Sqoop's half: a full-fidelity import of a slow-changing table."""
    db = os.path.join(tmp, "orders.db")
    con = sqlite3.connect(db)
    con.execute("""CREATE TABLE orders (
        order_id INTEGER PRIMARY KEY, host TEXT, region TEXT, revenue REAL)""")
    rows = []
    hosts = ["10.0.0.1", "10.0.0.2", "10.0.0.3", "10.0.0.4"]
    for i, (_, r) in enumerate(f.SALES_DF.iterrows()):
        rows.append((i + 1, hosts[i % 4], r["region"], float(r["revenue"])))
    con.executemany("INSERT INTO orders VALUES (?,?,?,?)", rows)
    con.commit()
    data = con.execute("SELECT order_id, host, region, revenue FROM orders").fetchall()
    con.close()
    path = os.path.join(tmp, "batch.parquet")
    pq.write_table(pa.Table.from_pylist(
        [dict(zip(("order_id", "host", "region", "revenue"), r)) for r in data]),
        path)
    return path, len(data), sum(r[3] for r in data)


def stream_leg(tmp, n):
    """Flume's half: events, parsed, with headers, landed as Parquet."""
    events = []
    for line in f.access_logs(n):
        host = line.split(" ", 1)[0]
        status = line.rsplit(" ", 2)[-2]
        size = int(line.rsplit(" ", 1)[-1])
        events.append({"host": host, "status": status, "bytes": size})
    path = os.path.join(tmp, "stream.parquet")
    pq.write_table(pa.Table.from_pylist(events), path)
    return path, len(events)


def main():
    print("  Experiment 14 -- batch and streaming, joined")

    # Step 1: Run the batch leg and the stream leg
    tmp = tempfile.mkdtemp(prefix="bigdata14_")
    batch, n_batch, batch_rev = batch_leg(tmp)
    stream, n_stream = stream_leg(tmp, 40)

    print(f"\n    batch  leg (Sqoop) : {n_batch} orders, "
          f"revenue {batch_rev:,.0f}")
    print(f"    stream leg (Flume) : {n_stream} events")
    assert abs(batch_rev - f.total_revenue()) < 1e-6

    con = duckdb.connect()
    con.execute(f"CREATE VIEW batch AS SELECT * FROM '{batch}'")
    con.execute(f"CREATE VIEW stream AS SELECT * FROM '{stream}'")

    # Step 2: Join them
    rows = con.execute("""
        SELECT b.host,
               COUNT(DISTINCT b.order_id) AS orders,
               SUM(DISTINCT b.revenue)    AS revenue,
               COUNT(s.host)              AS events,
               SUM(CASE WHEN s.status = '500' THEN 1 ELSE 0 END) AS errors
        FROM batch b LEFT JOIN stream s ON b.host = s.host
        GROUP BY b.host ORDER BY b.host
    """).fetchall()
    print(f"\n    joined on host:")
    print(f"      {'host':<12}{'orders':>8}{'events':>8}{'errors':>8}")
    for host, orders, rev, events, errors in rows:
        print(f"      {host:<12}{orders:>8}{events:>8}{errors:>8}")
    total_events = sum(r[3] for r in rows)
    assert total_events != n_stream, "the join FANS OUT -- see below"
    print(f"""
      events counted through the join: {total_events}
      events actually ingested       : {n_stream}""")
    print("""         THE JOIN INFLATED THE EVENT COUNT. Each host appears in
         several orders, so every event is counted once per matching
         order -- a FAN TRAP, and the same defect Course 11 found in
         a Power BI model. It is not a Spark problem, a Hive problem
         or a SQL problem; it is a GRAIN problem, and it appears
         wherever two fact tables are joined directly""")

    # Step 3: Fix the join
    fixed = con.execute("""
        WITH ev AS (
            SELECT host, COUNT(*) AS events,
                   SUM(CASE WHEN status = '500' THEN 1 ELSE 0 END) AS errors
            FROM stream GROUP BY host),
             ord AS (
            SELECT host, COUNT(*) AS orders, SUM(revenue) AS revenue
            FROM batch GROUP BY host)
        SELECT o.host, o.orders, o.revenue, e.events, e.errors
        FROM ord o JOIN ev e ON o.host = e.host ORDER BY o.host
    """).fetchall()
    print(f"\n    the fix -- aggregate EACH SIDE to a common grain FIRST:")
    print(f"      {'host':<12}{'orders':>8}{'revenue':>10}{'events':>8}{'errors':>8}")
    for host, orders, rev, events, errors in fixed:
        print(f"      {host:<12}{orders:>8}{rev:>10,.0f}{events:>8}{errors:>8}")
    assert sum(r[3] for r in fixed) == n_stream
    assert abs(sum(r[2] for r in fixed) - batch_rev) < 1e-6
    print(f"""         {sum(r[3] for r in fixed)} events and {sum(r[2] for r in fixed):,.0f} revenue -- both totals now
         reconcile with the sources. Aggregate to a shared grain, THEN
         join. That single rule prevents most wrong numbers in a data
         warehouse, and it is worth stating in exactly those words""")

    # Step 4: Compare the two legs
    print("\n    the two legs are not interchangeable:")
    print(f"      {'':<20}{'batch (Sqoop)':<26}{'streaming (Flume)'}")
    for label, b, s in (
            ("arrives", "on a schedule", "continuously"),
            ("completeness", "a whole table, consistent", "whatever has landed"),
            ("late data", "impossible", "NORMAL -- and must be handled"),
            ("re-runnable", "yes, idempotent", "no -- events are consumed"),
            ("catches DELETEs", "on a full re-import", "never"),
            ("file sizes", "large, controllable", "small unless you roll"),
            ("failure means", "re-run the import", "gap in the data")):
        print(f"      {label:<20}{b:<26}{s}")
    print("""         'late data is normal' is the row that changes the design.
         A streaming aggregate for 09:00 is not final at 10:00, so
         either you accept eventual correctness or you keep a
         watermark and re-emit. The batch leg has no such problem,
         which is why the LAMBDA ARCHITECTURE keeps both""")

    # Step 5: Set lambda against kappa
    print("\n    the two architectures this experiment is really about:")
    print(f"      {'':<10}{'layers':<34}{'cost'}")
    print(f"      {'Lambda':<10}{'batch + speed + serving':<34}"
          f"{'the logic is written TWICE'}")
    print(f"      {'Kappa':<10}{'one streaming path, replayable':<34}"
          f"{'needs a log like Kafka'}")
    print("""         Lambda's real cost is not machines, it is that the same
         business rule exists in two codebases and they drift. Kappa
         removes the batch layer by making the stream replayable --
         which is why Kafka replaced Flume in most of these pipelines
         after about 2016. Say that and you have placed the whole
         syllabus in time""")

    # Step 6: Reconcile
    status_counts = Counter(
        r[0] for r in con.execute("SELECT status FROM stream").fetchall())
    print(f"\n    reconciliation: {dict(sorted(status_counts.items()))}")
    assert status_counts["200"] == 24
    assert sum(status_counts.values()) == 40
    print("""         the same 24 / 8 / 8 as experiments 12 and 17. Three
         experiments, three code paths, one set of numbers -- and if
         a change ever breaks one of them the suite fails on all
         three, which is the only reason to build the check""")

    con.close()
    for path in (batch, stream, os.path.join(tmp, "orders.db")):
        os.remove(path)
    os.rmdir(tmp)


if __name__ == "__main__":
    main()

5. Execution and Results

The Python check, 14_pipeline.py:

OUTPUT

  Experiment 14 -- batch and streaming, joined

    batch  leg (Sqoop) : 9 orders, revenue 12,880
    stream leg (Flume) : 40 events

    joined on host:
      host          orders  events  errors
      10.0.0.1           3      30       6
      10.0.0.2           2      20       4
      10.0.0.3           2      20       4
      10.0.0.4           2      20       4

      events counted through the join: 90
      events actually ingested       : 40
         THE JOIN INFLATED THE EVENT COUNT. Each host appears in
         several orders, so every event is counted once per matching
         order -- a FAN TRAP, and the same defect Course 11 found in
         a Power BI model. It is not a Spark problem, a Hive problem
         or a SQL problem; it is a GRAIN problem, and it appears
         wherever two fact tables are joined directly

    the fix -- aggregate EACH SIDE to a common grain FIRST:
      host          orders   revenue  events  errors
      10.0.0.1           3     4,200      10       2
      10.0.0.2           2     3,220      10       2
      10.0.0.3           2     2,800      10       2
      10.0.0.4           2     2,660      10       2
         40 events and 12,880 revenue -- both totals now
         reconcile with the sources. Aggregate to a shared grain, THEN
         join. That single rule prevents most wrong numbers in a data
         warehouse, and it is worth stating in exactly those words

    the two legs are not interchangeable:
                          batch (Sqoop)             streaming (Flume)
      arrives             on a schedule             continuously
      completeness        a whole table, consistent whatever has landed
      late data           impossible                NORMAL -- and must be handled
      re-runnable         yes, idempotent           no -- events are consumed
      catches DELETEs     on a full re-import       never
      file sizes          large, controllable       small unless you roll
      failure means       re-run the import         gap in the data
         'late data is normal' is the row that changes the design.
         A streaming aggregate for 09:00 is not final at 10:00, so
         either you accept eventual correctness or you keep a
         watermark and re-emit. The batch leg has no such problem,
         which is why the LAMBDA ARCHITECTURE keeps both

    the two architectures this experiment is really about:
                layers                            cost
      Lambda    batch + speed + serving           the logic is written TWICE
      Kappa     one streaming path, replayable    needs a log like Kafka
         Lambda's real cost is not machines, it is that the same
         business rule exists in two codebases and they drift. Kappa
         removes the batch layer by making the stream replayable --
         which is why Kafka replaced Flume in most of these pipelines
         after about 2016. Say that and you have placed the whole
         syllabus in time

    reconciliation: {'200': 24, '404': 8, '500': 8}
         the same 24 / 8 / 8 as experiments 12 and 17. Three
         experiments, three code paths, one set of numbers -- and if
         a change ever breaks one of them the suite fails on all
         three, which is the only reason to build the check

THE FIX

host orders revenue events errors
10.0.0.1 3 4,200 10 2
10.0.0.2 2 3,220 10 2
10.0.0.3 2 2,800 10 2
10.0.0.4 2 2,660 10 2
total 9 12,880 40 8

Aggregate each side to a common grain first, then join. Both totals now reconcile with the sources. That single rule prevents most wrong numbers in a data warehouse.

The row that changes the design. "Late data is normal." A streaming aggregate for 09:00 is not final at 10:00, so either you accept eventual correctness or you keep a watermark and re-emit. The batch leg has no such problem — which is exactly why the Lambda architecture keeps both, and why Kappa removes the batch layer by making the stream replayable.

Lambda's real cost is not machines — it is the same business rule living in two codebases and drifting.

RESULT

Joined directly, the 40 events are counted 90 times — the fan trap. Aggregated to the host first, then joined, both sides reconcile: 9 orders, ₹12,880, 40 events, 8 errors.

Experiment 15 — HBase: tables and CRUD

1. Question

Create and manage tables in HBase, with CRUD operations.

2. Aim

Create a table with two column families, put, get and scan cells, update and delete them, count atomically, alter and truncate the table and pre-split another; then model row keys, versions and tombstones.

3. Steps

On the cluster, 15_hbase.rb:

  1. Create the table.
  2. Put cells.
  3. Get a row.
  4. Scan.
  5. Update and delete.
  6. Count atomically.
  7. Alter, and truncate.
  8. Pre-split a table.

The Python check, 15_hbase_model.py:

  1. Create the table and put rows.
  2. Get a row.
  3. Keep versions.
  4. Delete, with a tombstone.
  5. Scan by prefix.
  6. Design the row key.
  7. Compare HBase with what it is confused with.

THE ROW KEY THAT SILENTLY LOSES A SALE

row key 'region#store#date' over 9 fact rows
produces only 8 DISTINCT KEYS -- 1 row would be overwritten

Vijayawada sold Rice and Shampoo on D1. HBase would not complain — it would version one over the other. A row key must be unique at the grain, and in HBase nothing checks that for you.

4. Programme

On the cluster, 15_hbase.rb:

# Experiment 15 -- create and manage tables in HBase -- CRUD operations
#
# Run it: hbase shell 15_hbase.rb, with HBase running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 15_hbase_model.py, which implements the same data model and runs it
#
# run with:  hbase shell 15_hbase.rb        (or paste into an interactive shell)

# Step 1: Create the table
create 'sales', \
  {NAME => 'info',  VERSIONS => 3, COMPRESSION => 'SNAPPY'}, \
  {NAME => 'sales', VERSIONS => 3, TTL => 31536000}
#   COLUMN FAMILIES ARE FIXED AT CREATE TIME and expensive to change.
#   COLUMNS inside a family are free and need no declaration -- that is what
#   "schema-less" means in HBase, and it is only half true.
#   Keep families to two or three: each one is a separate store file, and a
#   flush of one flushes all of them.

list
describe 'sales'

# Step 2: Put cells
# row key = region#store#date#product -- composite, unique at the grain,
# and NOT monotonic. See the design table at the end.
put 'sales', 'South#Vijayawada#D1#Rice',  'info:product',  'Rice 5kg'
put 'sales', 'South#Vijayawada#D1#Rice',  'info:category', 'Grocery'
put 'sales', 'South#Vijayawada#D1#Rice',  'sales:qty',     '10'
put 'sales', 'South#Vijayawada#D1#Rice',  'sales:revenue', '2800'
put 'sales', 'North#Hyderabad#D2#Notebook', 'info:product', 'Notebook'
put 'sales', 'North#Hyderabad#D2#Notebook', 'sales:qty',    '20'

# Step 3: Get a row
get 'sales', 'South#Vijayawada#D1#Rice'
get 'sales', 'South#Vijayawada#D1#Rice', 'sales'
get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
#   a PUT to an existing cell ADDS A VERSION; it does not overwrite.

# Step 4: Scan
scan 'sales'
scan 'sales', {LIMIT => 5}
scan 'sales', {STARTROW => 'South', STOPROW => 'South~'}
#   '~' sorts after every printable ASCII letter, which is the idiomatic way
#   to write a prefix scan by hand. PrefixFilter does the same thing:
scan 'sales', {FILTER => "PrefixFilter('South')"}
scan 'sales', {FILTER => "SingleColumnValueFilter('info','category',=,'binary:Grocery',true,true)"}
#   [Corrected: without the two trues -- filterIfMissing and latestVersionOnly --
#   a row that has NO info:category passes the filter, and the scan returned the
#   Notebook row too. A filter on a column says nothing about rows without it.]
#   ^ THIS IS A FULL TABLE SCAN. HBase has no secondary index. The filter runs
#     server-side, so less data crosses the network -- but every row is read.

# Step 5: Update and delete
put    'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '11'   # = a new version
get    'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
#   [Changed: this get is added. The one above ran before any second put, so it
#   showed one version; here are both, 11 over 10.]
delete 'sales', 'South#Vijayawada#D1#Rice', 'info:category'
deleteall 'sales', 'North#Hyderabad#D2#Notebook'
#   A DELETE WRITES A TOMBSTONE. The table gets BIGGER. Data and marker are
#   both removed only at a major compaction:
major_compact 'sales'

# Step 6: Count atomically
incr 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views', 1
get_counter 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views'

# Step 7: Alter, and truncate
count 'sales', INTERVAL => 100
disable 'sales'
alter 'sales', {NAME => 'info', VERSIONS => 5}
enable 'sales'
truncate 'sales'          # = disable + drop + recreate. Keeps the schema.
# drop 'sales'            # must be disabled first

# Step 8: Pre-split a table
create 'sales2', 'info', {SPLITS => ['East', 'North', 'South', 'West']}
#   without pre-splits the table starts as ONE region on ONE RegionServer, so
#   a bulk load runs single-threaded until the first split.

# --- row key design ---------------------------------------------------------
#   key                          writes go to     verdict
#   timestamp                    the last region  HOTSPOT
#   sequential id                the last region  HOTSPOT
#   md5(id) + id                 everywhere       good; range scans lost
#   region#store#date#product    by region        good; prefix scans work
#
#   You cannot have even write distribution AND range scans on the same
#   dimension. Choosing between them IS row-key design.

The Python check, 15_hbase_model.py:

"""Experiment 15 -- create and manage tables in HBase (CRUD operations).

`15_hbase.rb` carries the real shell commands, and runs in the HBase shell (the
lab page shows it). What runs here is HBase's DATA MODEL, implemented honestly:
a sorted map from (row, family:qualifier, version) to bytes, with real
versioning, real tombstones and real row-key range scans.

The model IS the exam. Almost every HBase question -- why scans are fast and
gets by value are not, why a monotonic row key is a disaster, why a delete
does not free space -- is a consequence of "sorted map, sharded by row-key
range".
"""
import bisect

import fixtures as f


class HBase:
    """(row, family, qualifier) -> {timestamp: value}, kept SORTED BY ROW."""

    def __init__(self, families, max_versions=3):
        self.families = set(families)
        self.max_versions = max_versions
        self.cells = {}
        self.rows = []              # sorted, because everything depends on it
        self.clock = 0

    def _tick(self):
        self.clock += 1
        return self.clock

    def put(self, row, fam, qual, value):
        if fam not in self.families:
            raise KeyError(f"column family {fam!r} was not declared at create time")
        if row not in self.cells:
            bisect.insort(self.rows, row)
            self.cells[row] = {}
        versions = self.cells[row].setdefault((fam, qual), {})
        versions[self._tick()] = value
        for ts in sorted(versions)[:-self.max_versions]:
            del versions[ts]

    def get(self, row, fam=None, qual=None, versions=1):
        if row not in self.cells:
            return {}
        out = {}
        for (fm, q), vs in self.cells[row].items():
            if fam and fm != fam:
                continue
            if qual and q != qual:
                continue
            live = []
            for ts, v in sorted(vs.items(), reverse=True):
                if v is None:
                    break          # a tombstone MASKS every older version
                live.append((ts, v))
            live = live[:versions]
            if live:
                out[f"{fm}:{q}"] = live if versions > 1 else live[0][1]
        return out

    def delete(self, row, fam, qual):
        """A delete writes a TOMBSTONE. It does not remove anything."""
        self.cells[row].setdefault((fam, qual), {})[self._tick()] = None

    def scan(self, start=None, stop=None):
        lo = bisect.bisect_left(self.rows, start) if start else 0
        hi = bisect.bisect_left(self.rows, stop) if stop else len(self.rows)
        return [(r, self.get(r)) for r in self.rows[lo:hi]]

    def storefiles(self):
        """Every version and every tombstone still occupies a cell."""
        return sum(len(v) for row in self.cells.values() for v in row.values())


def main():
    print("  Experiment 15 -- the HBase data model, implemented")

    t = HBase(families={"info", "sales"}, max_versions=3)

    # Step 1: Create the table and put rows
    # First, a row key that looks reasonable and is NOT unique at the grain.
    naive = {f"{r['region']}#{r['store']}#{r['date_key']}"
             for _, r in f.SALES_DF.iterrows()}
    print(f"\n    row key 'region#store#date' over {len(f.SALES_DF)} fact rows")
    print(f"    produces only {len(naive)} DISTINCT KEYS -- "
          f"{len(f.SALES_DF) - len(naive)} row would be overwritten")
    assert len(naive) == 8, "two facts share a store and a date"
    print("""         Vijayawada sold Rice AND Shampoo on D1, so those two
         facts collide. HBase would not complain -- it would simply
         version one over the other and lose a sale.
         A row key must be UNIQUE AT THE GRAIN. In an RDBMS the
         primary key declaration catches this; in HBase nothing does,
         and that is the failure mode to remember""")

    for _, r in f.SALES_DF.iterrows():
        # row key: region#store#date#product -- unique, composite, NOT monotonic
        key = (f"{r['region']}#{r['store']}#{r['date_key']}#"
               f"{r['product'].split()[0]}")
        t.put(key, "info", "product", r["product"])
        t.put(key, "info", "category", r["category"])
        t.put(key, "sales", "qty", int(r["qty"]))
        t.put(key, "sales", "revenue", float(r["revenue"]))

    print(f"\n    {len(t.rows)} rows, {t.storefiles()} cells")
    print(f"    row keys are SORTED, always:")
    for k in t.rows[:4]:
        print(f"      {k}")
    print(f"      ... {len(t.rows) - 4} more")
    assert t.rows == sorted(t.rows)

    # Step 2: Get a row
    key = t.rows[0]
    print(f"\n    GET '{key}':")
    for col, val in sorted(t.get(key).items()):
        print(f"      {col:<18}{val}")

    # Step 3: Keep versions
    t.put(key, "sales", "qty", 99)
    t.put(key, "sales", "qty", 111)
    vs = t.get(key, "sales", "qty", versions=3)["sales:qty"]
    print(f"\n    after two more PUTs to the same cell, 3 versions:")
    for ts, v in vs:
        print(f"      ts={ts:<5}{v}")
    assert [v for _, v in vs][:2] == [111, 99]
    assert len(vs) == 3, "VERSIONS => 3 caps the history at three"
    print("""         a PUT to an existing cell does not overwrite -- it adds a
         VERSION, and the old value is still readable. VERSIONS => 3
         at create time is what caps it. That is why HBase is
         described as a multidimensional map: row, family, qualifier
         AND time""")

    # Step 4: Delete, with a tombstone
    before = t.storefiles()
    t.delete(key, "info", "category")
    after = t.storefiles()
    assert "info:category" not in t.get(key)
    assert after > before, "a delete makes the table BIGGER until compaction"
    print(f"\n    DELETE info:category")
    print(f"      readable? {'yes' if 'info:category' in t.get(key) else 'no'}")
    print(f"      cells:    {before} -> {after}")
    print("""         THE TABLE GOT BIGGER. A delete writes a tombstone marker;
         the data and the marker both disappear only at MAJOR
         COMPACTION. This is the answer to 'I deleted a billion rows
         and disk usage went up'""")

    # Step 5: Scan by prefix
    south = t.scan("South", "South~")
    north = t.scan("North", "North~")
    print(f"\n    SCAN 'South' .. 'South~' -> {len(south)} rows")
    print(f"    SCAN 'North' .. 'North~' -> {len(north)} rows")
    assert len(south) + len(north) == len(t.rows)
    assert len(south) == 6 and len(north) == 3
    assert len(t.rows) == 9, "the unique key keeps all nine facts"
    print("""         a range scan on the row-key PREFIX reads exactly the
         rows you want, sequentially, from one or two regions. That
         is the fastest thing HBase does -- and it works only because
         'region' is the FIRST component of the key""")

    print("\n    the same question asked the wrong way:")
    matches = [r for r, cols in t.scan() if cols.get("info:category") == "Grocery"]
    print(f"      find category = 'Grocery' -> {len(matches)} rows, "
          f"after scanning all {len(t.rows)}")
    print("""         HBase has NO SECONDARY INDEX. Filtering on a value means
         a FULL TABLE SCAN with a server-side filter -- correct, and
         O(table). If you need that query, you build a second table
         keyed by category, and you keep it in sync yourself""")

    # Step 6: Design the row key
    print("\n    row key design, which is the whole job:")
    print(f"      {'key':<34}{'regions hit by a write':<24}verdict")
    for key_desc, hits, verdict in (
            ("timestamp (1723459200, ...)", "ONE -- always the last", "HOTSPOT"),
            ("sequential id (1, 2, 3, ...)", "ONE -- always the last", "HOTSPOT"),
            ("md5(id) + id", "all, evenly", "good, scans lost"),
            ("region#store#date", "by region", "good, prefix scans work")):
        print(f"      {key_desc:<34}{hits:<24}{verdict}")
    print("""         a monotonically increasing row key sends EVERY write to
         the same RegionServer, so a 50-node cluster runs at the
         speed of one node. Salting or hashing fixes the hotspot and
         destroys range scans -- you cannot have both, and choosing
         is what row-key design means""")

    # Step 7: Compare HBase with what it is confused with
    print("\n    HBase against what students compare it to:")
    print(f"      {'':<14}{'HBase':<26}{'Hive':<22}{'MongoDB (Course 10)'}")
    for label, hb, hv, mg in (
            ("model", "sparse sorted map", "tables over files", "documents"),
            ("latency", "milliseconds", "seconds to minutes", "milliseconds"),
            ("random writes", "YES", "no", "YES"),
            ("secondary index", "no", "no", "YES"),
            ("query language", "get/put/scan only", "HiveQL", "MQL"),
            ("schema", "families fixed, cols free", "fixed", "free")):
        print(f"      {label:<14}{hb:<26}{hv:<22}{mg}")
    print("""         HBase and Hive both sit on HDFS and answer completely
         different questions: Hive scans everything slowly, HBase
         fetches one row instantly. And note the row students always
         get wrong -- HBase has NO secondary index where MongoDB
         does, which is the sharpest difference between the two
         NoSQL stores this programme teaches""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 15_hbase.rb:

OUTPUT

$ start-hbase.sh > /dev/null 2>&1
$ sleep 20
$ hbase shell -n < 15_hbase.typed 2>&1
hbase:001:0> create 'sales', \
hbase:002:0*   {NAME => 'info',  VERSIONS => 3, COMPRESSION => 'SNAPPY'}, \
hbase:003:0*   {NAME => 'sales', VERSIONS => 3, TTL => 31536000}
Created table sales
Took 1.2161 seconds
=> Hbase::Table - sales
hbase:004:0> list
TABLE
sales
1 row(s)
Took 0.0266 seconds
=> ["sales"]
hbase:005:0> describe 'sales'
Table sales is ENABLED
sales, {TABLE_ATTRIBUTES => {METADATA => {'hbase.store.file-tracker.impl' => 'DEFAULT'}}}
COLUMN FAMILIES DESCRIPTION
{NAME => 'info', INDEX_BLOCK_ENCODING => 'NONE', VERSIONS => '3', KEEP_DELETED_CELLS => 'FALSE', DATA_BLOCK_ENCODING => 'NONE', TTL => 'FOREVER', MIN_VERSIONS => '0', REPLICATION_SCOPE => '0', BLOOMFILTER => 'ROW', IN_MEMORY => 'false', COMPRESSION => 'SNAPPY', BLOCKCACHE => 'true', BLOCKSIZE => '65536 B (64KB)'}

{NAME => 'sales', INDEX_BLOCK_ENCODING => 'NONE', VERSIONS => '3', KEEP_DELETED_CELLS => 'FALSE', DATA_BLOCK_ENCODING => 'NONE', TTL => '31536000 SECONDS (365 DAYS)', MIN_VERSIONS => '0', REPLICATION_SCOPE => '0', BLOOMFILTER => 'ROW', IN_MEMORY => 'false', COMPRESSION => 'NONE', BLOCKCACHE => 'true', BLOCKSIZE => '65536 B (64KB)'}

2 row(s)
Quota is disabled
Took 0.1691 seconds
hbase:006:0> put 'sales', 'South#Vijayawada#D1#Rice',  'info:product',  'Rice 5kg'
Took 0.0882 seconds
hbase:007:0> put 'sales', 'South#Vijayawada#D1#Rice',  'info:category', 'Grocery'
Took 0.0052 seconds
hbase:008:0> put 'sales', 'South#Vijayawada#D1#Rice',  'sales:qty',     '10'
Took 0.0083 seconds
hbase:009:0> put 'sales', 'South#Vijayawada#D1#Rice',  'sales:revenue', '2800'
Took 0.0065 seconds
hbase:010:0> put 'sales', 'North#Hyderabad#D2#Notebook', 'info:product', 'Notebook'
Took 0.0051 seconds
hbase:011:0> put 'sales', 'North#Hyderabad#D2#Notebook', 'sales:qty',    '20'
Took 0.0060 seconds
hbase:012:0> get 'sales', 'South#Vijayawada#D1#Rice'
COLUMN  CELL
 info:category timestamp=2026-10-04T23:30:58.482, value=Grocery
 info:product timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 sales:qty timestamp=2026-10-04T23:30:58.496, value=10
 sales:revenue timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0716 seconds
hbase:013:0> get 'sales', 'South#Vijayawada#D1#Rice', 'sales'
COLUMN  CELL
 sales:qty timestamp=2026-10-04T23:30:58.496, value=10
 sales:revenue timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0102 seconds
hbase:014:0> get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
COLUMN  CELL
 sales:qty timestamp=2026-10-04T23:30:58.496, value=10
1 row(s)
Took 0.0093 seconds
hbase:015:0> scan 'sales'
ROW  COLUMN+CELL
 North#Hyderabad#D2#Notebook column=info:product, timestamp=2026-10-04T23:30:58.520, value=Notebook
 North#Hyderabad#D2#Notebook column=sales:qty, timestamp=2026-10-04T23:30:58.532, value=20
 South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
 South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
 South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
2 row(s)
Took 0.0168 seconds
hbase:016:0> scan 'sales', {LIMIT => 5}
ROW  COLUMN+CELL
 North#Hyderabad#D2#Notebook column=info:product, timestamp=2026-10-04T23:30:58.520, value=Notebook
 North#Hyderabad#D2#Notebook column=sales:qty, timestamp=2026-10-04T23:30:58.532, value=20
 South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
 South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
 South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
2 row(s)
Took 0.0306 seconds
hbase:017:0> scan 'sales', {STARTROW => 'South', STOPROW => 'South~'}
ROW  COLUMN+CELL
 South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
 South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
 South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0142 seconds
hbase:018:0> scan 'sales', {FILTER => "PrefixFilter('South')"}
ROW  COLUMN+CELL
 South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
 South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
 South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0252 seconds
hbase:019:0> scan 'sales', {FILTER => "SingleColumnValueFilter('info','category',=,'binary:Grocery',true,true)"}
ROW  COLUMN+CELL
 South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
 South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
 South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
 South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0670 seconds
hbase:020:0> put    'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '11'   # = a new version
Took 0.0092 seconds
hbase:021:0> get    'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
COLUMN  CELL
 sales:qty timestamp=2026-10-04T23:30:58.881, value=11
 sales:qty timestamp=2026-10-04T23:30:58.496, value=10
1 row(s)
Took 0.0083 seconds
hbase:022:0> delete 'sales', 'South#Vijayawada#D1#Rice', 'info:category'
Took 0.0119 seconds
hbase:023:0> deleteall 'sales', 'North#Hyderabad#D2#Notebook'
Took 0.0048 seconds
hbase:024:0> major_compact 'sales'
Took 0.0647 seconds
hbase:025:0> incr 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views', 1
COUNTER VALUE = 1
Took 0.0173 seconds
hbase:026:0> get_counter 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views'
COUNTER VALUE = 1
Took 0.0045 seconds
hbase:027:0> count 'sales', INTERVAL => 100
2 row(s)
Took 0.0157 seconds
=> 2
hbase:028:0> disable 'sales'
Took 0.7025 seconds
hbase:029:0> alter 'sales', {NAME => 'info', VERSIONS => 5}
Updating all regions with the new schema...
All regions updated.
Done.
Took 1.2158 seconds
hbase:030:0> enable 'sales'
Took 0.6759 seconds
hbase:031:0> truncate 'sales'          # = disable + drop + recreate. Keeps the schema.
Truncating 'sales' table (it may take a while):
Disabling table...
Truncating table...
Took 0.9969 seconds
hbase:032:0> create 'sales2', 'info', {SPLITS => ['East', 'North', 'South', 'West']}
Created table sales2
Took 0.6325 seconds
=> Hbase::Table - sales2
hbase:033:0>
$ stop-hbase.sh 2>&1 | tail -1
stopping hbase..............

The Python check, 15_hbase_model.py:

OUTPUT

  Experiment 15 -- the HBase data model, implemented

    row key 'region#store#date' over 9 fact rows
    produces only 8 DISTINCT KEYS -- 1 row would be overwritten
         Vijayawada sold Rice AND Shampoo on D1, so those two
         facts collide. HBase would not complain -- it would simply
         version one over the other and lose a sale.
         A row key must be UNIQUE AT THE GRAIN. In an RDBMS the
         primary key declaration catches this; in HBase nothing does,
         and that is the failure mode to remember

    9 rows, 36 cells
    row keys are SORTED, always:
      North#Hyderabad#D2#Notebook
      North#Hyderabad#D3#Rice
      North#Hyderabad#D4#Notebook
      South#Guntur#D1#Tea
      ... 5 more

    GET 'North#Hyderabad#D2#Notebook':
      info:category     Stationery
      info:product      Notebook
      sales:qty         20
      sales:revenue     800.0

    after two more PUTs to the same cell, 3 versions:
      ts=38   111
      ts=37   99
      ts=19   20
         a PUT to an existing cell does not overwrite -- it adds a
         VERSION, and the old value is still readable. VERSIONS => 3
         at create time is what caps it. That is why HBase is
         described as a multidimensional map: row, family, qualifier
         AND time

    DELETE info:category
      readable? no
      cells:    38 -> 39
         THE TABLE GOT BIGGER. A delete writes a tombstone marker;
         the data and the marker both disappear only at MAJOR
         COMPACTION. This is the answer to 'I deleted a billion rows
         and disk usage went up'

    SCAN 'South' .. 'South~' -> 6 rows
    SCAN 'North' .. 'North~' -> 3 rows
         a range scan on the row-key PREFIX reads exactly the
         rows you want, sequentially, from one or two regions. That
         is the fastest thing HBase does -- and it works only because
         'region' is the FIRST component of the key

    the same question asked the wrong way:
      find category = 'Grocery' -> 5 rows, after scanning all 9
         HBase has NO SECONDARY INDEX. Filtering on a value means
         a FULL TABLE SCAN with a server-side filter -- correct, and
         O(table). If you need that query, you build a second table
         keyed by category, and you keep it in sync yourself

    row key design, which is the whole job:
      key                               regions hit by a write  verdict
      timestamp (1723459200, ...)       ONE -- always the last  HOTSPOT
      sequential id (1, 2, 3, ...)      ONE -- always the last  HOTSPOT
      md5(id) + id                      all, evenly             good, scans lost
      region#store#date                 by region               good, prefix scans work
         a monotonically increasing row key sends EVERY write to
         the same RegionServer, so a 50-node cluster runs at the
         speed of one node. Salting or hashing fixes the hotspot and
         destroys range scans -- you cannot have both, and choosing
         is what row-key design means

    HBase against what students compare it to:
                    HBase                     Hive                  MongoDB (Course 10)
      model         sparse sorted map         tables over files     documents
      latency       milliseconds              seconds to minutes    milliseconds
      random writes YES                       no                    YES
      secondary indexno                        no                    YES
      query languageget/put/scan only         HiveQL                MQL
      schema        families fixed, cols free fixed                 free
         HBase and Hive both sit on HDFS and answer completely
         different questions: Hive scans everything slowly, HBase
         fetches one row instantly. And note the row students always
         get wrong -- HBase has NO secondary index where MongoDB
         does, which is the sharpest difference between the two
         NoSQL stores this programme teaches

_drive_15_hbase.py starts HBase in standalone mode — its own ZooKeeper, data in a local folder — with the Java Snappy codec configured, since the table asks for COMPRESSION => 'SNAPPY' and the native library is not there; then types the file into hbase shell.

Versions. A put to an existing cell adds a version; it does not overwrite. The shell's get with VERSIONS => 3 above shows both values of sales:qty, newest first; the model:

ts=38   111
ts=37    99
ts=19    20

A DELETE MAKES THE TABLE BIGGER

DELETE info:category   readable? no    cells: 38 -> 39

A tombstone. Data and marker disappear only at major compaction — the answer to "I deleted a billion rows and disk usage went up".

Scans. SCAN 'South'..'South~' → 6 rows in the model's nine, and 'North'..'North~' → 3; in the shell's two-row table, 1. A prefix scan reads exactly the rows you want, sequentially — and works only because region is the first component of the key. Ask it the wrong way — category = 'Grocery' — and it is a full table scan. HBase has no secondary index.

ROW KEY DESIGN

Key Writes go to Verdict
timestamp the last region HOTSPOT
sequential id the last region HOTSPOT
md5(id) + id everywhere good — scans lost
region#store#date#product by region good — prefix scans work

You cannot have even write distribution and range scans on the same dimension. Choosing is what row-key design means.

RESULT

In the HBase shell every operation ran: a second put kept both versions of the cell, a prefix scan and a value filter each found the one South row, the counter counted to 1. The filter needed two more arguments, or rows without the column passed it. In the model, the row key region#store#date loses a sale.

Experiment 16 — Coordination with ZooKeeper

1. Question

Demonstrate coordination with ZooKeeper.

2. Aim

Start a three-server ensemble, see which server leads, build a tree of znodes, make ephemeral and sequential nodes, and run the leader-election recipe; then model election, locking and ensemble sizing.

3. Steps

On the cluster, 16_zookeeper.sh:

  1. Configure three servers.
  2. Start them, and see the leader.
  3. Build the tree.
  4. Make an ephemeral node.
  5. Make sequential nodes.
  6. Elect a leader.
  7. Look for the systems that use it.

The Python check, 16_zookeeper_model.py:

  1. Elect a leader.
  2. Take a distributed lock.
  3. Rely on an atomic create.
  4. Size the ensemble.
  5. See what ZooKeeper is not.
  6. Name who uses it.

LEADER ELECTION

nn1 -> lock-0000000000     LEADER: nn1
nn2 -> lock-0000000001
nn3 -> lock-0000000002

nn1's session expires:  new LEADER: nn2

Nobody ran a failover script. The ephemeral node was deleted by the server, the watch fired, and nn2 saw itself at the head of the queue. That is how HDFS NameNode HA works — the link back to experiment 5.

4. Programme

On the cluster, 16_zookeeper.sh:

# Experiment 16 -- demonstrate coordination with ZooKeeper
#
# Run it: bash 16_zookeeper.sh -- it starts and stops its own ensemble. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 16_zookeeper_model.py, which runs leader election, locking and quorum maths
#
# Step 1: Configure three servers
# conf/zoo.cfg, identical on all three:
#   tickTime=2000
#   initLimit=10
#   syncLimit=5
#   dataDir=/var/lib/zookeeper
#   clientPort=2181
#   server.1=zk1:2888:3888
#   server.2=zk2:2888:3888
#   server.3=zk3:2888:3888
# and on each host:  echo <id> > /var/lib/zookeeper/myid
#
# THREE, FIVE OR SEVEN. Four servers need a quorum of 3 and tolerate one
# failure -- exactly what three tolerate -- so the fourth machine buys nothing
# and slows every write.
#
# Here the three servers run on ONE machine, so each needs its own client port,
# its own pair of quorum ports and its own data directory; otherwise the file
# is the one above. The four-letter commands must also be allowed by name:
for id in 1 2 3; do
  mkdir -p zk$id && echo $id > zk$id/myid
  cat > zk$id.cfg <<EOF
tickTime=2000
initLimit=10
syncLimit=5
dataDir=$PWD/zk$id
clientPort=218$id
server.1=localhost:2888:3888
server.2=localhost:2889:3889
server.3=localhost:2890:3890
4lw.commands.whitelist=srvr,mntr
EOF
done
# [Corrected: the ensemble was described in comments, and `zkServer.sh start`
# started ONE server. These are the zoo.cfg above for three servers on one
# host, and each is started below.]

# Step 2: Start them, and see the leader
for id in 1 2 3; do ZOO_LOG_DIR=zk$id zkServer.sh start $PWD/zk$id.cfg 2>&1 | tail -1; done
sleep 10                            # they elect a leader
for id in 1 2 3; do zkServer.sh status $PWD/zk$id.cfg 2>&1 | grep Mode; done
#                                     "Mode: leader" on exactly one server
echo srvr | nc localhost 2181 | grep -E "Mode|Zxid|Node count"     # state and zxid
echo mntr | nc localhost 2181 | grep -E "zk_server_state|zk_znode_count|zk_outstanding_requests"

# Step 3: Build the tree
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /
create /app "config-v1"
get /app
set /app "config-v2"
stat /app
ls -R /
quit
EOF
#   stat: dataVersion increments; cZxid, mZxid

# Step 4: Make an ephemeral node
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -e /app/worker-1 "alive"
ls /app
quit
EOF
# reconnect: the session ended, so the ephemeral node is GONE
zkCli.sh -server localhost:2182 2>&1 <<'EOF'
ls /app
quit
EOF
#   (asked of a different server: every one of them has the same tree)

# Step 5: Make sequential nodes
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -s /app/task- "t"
create -s /app/task- "t"
create -s /app/task- "t"
quit
EOF
#   -> /app/task-0000000001, -0000000002, -0000000003
# [Corrected: this said -0000000000, -0000000001, -0000000002. The number is
# kept by the PARENT, /app, and counts every child it has had: worker-1 above
# took the first, though it is gone. Never assume a sequence starts at 0.]

# Step 6: Elect a leader
#   1. every candidate: create -e -s /election/n-
#   2. read the children; LOWEST sequence number is the leader
#   3. everyone else watches the node IMMEDIATELY BELOW their own
#      -- not the leader. Watching the leader wakes every candidate on one
#         failure: the HERD EFFECT.
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create /election ""
create -e -s /election/n- "nn1"
create -e -s /election/n- "nn2"
ls /election
get -w /election/n-0000000000
quit
EOF
#   get -w sets a ONE-SHOT watch

# Step 7: Look for the systems that use it
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /hbase
ls /hadoop-ha/mycluster
ls /rmstore
quit
EOF
#   /hbase                 master, rs, meta-region-server
#   /hadoop-ha/mycluster   ActiveStandbyElectorLock  <- NameNode HA
#   /rmstore               YARN ResourceManager HA
# Nothing uses this ensemble, so none of them exists here. Every HA story in
# the Hadoop ecosystem ends in znodes like these.

zkCli.sh -server localhost:2181 deleteall /app 2>/dev/null | tail -1
for id in 1 2 3; do zkServer.sh stop $PWD/zk$id.cfg 2>&1 | tail -1; done
# [Corrected: the zkCli.sh commands were indented under `zkCli.sh -server ...`,
# as typed at its prompt. Run as a script, the shell runs them itself -- `ls /`
# lists the machine's root directory, and `create` is "command not found".
# Each session is a here-document here, typed into zkCli.sh.]

The Python check, 16_zookeeper_model.py:

"""Experiment 16 -- demonstrate coordination with ZooKeeper.

`16_zookeeper.sh` carries the real `zkCli.sh` sessions, and runs on a real
three-server ensemble (the lab page shows it). What runs here is the
COORDINATION LOGIC: a
znode tree with ephemeral and sequential nodes, leader election by the
standard recipe, a distributed lock, and the quorum arithmetic that decides
whether an ensemble can make progress at all.

The point: ZooKeeper is not a database and not a queue. It is a small,
strongly-consistent tree whose only interesting properties are (a) ephemeral
nodes vanish when a session dies and (b) sequential nodes are numbered by a
single authority. Every recipe is built from exactly those two facts.
"""


class ZooKeeper:
    def __init__(self):
        self.tree = {"/": {"data": None, "ephemeral": False, "children": []}}
        self.counter = {}
        self.sessions = {}
        self.watches = []

    def create(self, path, data=None, ephemeral=False, sequential=False,
               session=None):
        parent = path.rsplit("/", 1)[0] or "/"
        if parent not in self.tree:
            raise KeyError(f"no node {parent} -- ZooKeeper creates no parents")
        if sequential:
            n = self.counter.get(path, 0)
            self.counter[path] = n + 1
            path = f"{path}{n:010d}"
        if path in self.tree:
            raise FileExistsError(f"{path} exists -- create is ATOMIC")
        self.tree[path] = {"data": data, "ephemeral": ephemeral,
                           "children": [], "session": session}
        self.tree[parent]["children"].append(path)
        if ephemeral:
            self.sessions.setdefault(session, []).append(path)
        return path

    def children(self, path):
        return sorted(self.tree[path]["children"])

    def expire(self, session):
        """A session dies -- every ephemeral node it owns disappears."""
        gone = self.sessions.pop(session, [])
        for p in gone:
            parent = p.rsplit("/", 1)[0] or "/"
            self.tree[parent]["children"].remove(p)
            del self.tree[p]
            self.watches.append(("deleted", p))
        return gone


def quorum(n):
    return n // 2 + 1


def main():
    print("  Experiment 16 -- ZooKeeper coordination")

    zk = ZooKeeper()
    zk.create("/hadoop-ha")
    zk.create("/hadoop-ha/mycluster")

    # Step 1: Elect a leader
    print("\n    leader election, the standard recipe:")
    print("      every candidate creates an EPHEMERAL SEQUENTIAL znode")
    print("      the LOWEST sequence number is the leader")
    print("      everyone else watches the node just below them")
    nodes = {}
    for host in ("nn1", "nn2", "nn3"):
        p = zk.create("/hadoop-ha/mycluster/lock-", data=host,
                      ephemeral=True, sequential=True, session=host)
        nodes[host] = p
        print(f"      {host} -> {p.rsplit('/', 1)[1]}")
    order = zk.children("/hadoop-ha/mycluster")
    leader = zk.tree[order[0]]["data"]
    print(f"\n      LEADER: {leader}")
    assert leader == "nn1"

    print("\n      nn1's session expires (its JVM was killed):")
    gone = zk.expire("nn1")
    order = zk.children("/hadoop-ha/mycluster")
    new_leader = zk.tree[order[0]]["data"]
    print(f"      {gone[0].rsplit('/', 1)[1]} vanished; new LEADER: {new_leader}")
    assert new_leader == "nn2" and len(order) == 2
    print("""         nobody ran a failover script. The ephemeral node was
         deleted BY THE SERVER when the heartbeat stopped, the watch
         fired, and nn2 saw itself at the head of the queue. That is
         how HDFS NameNode HA actually chooses its active node --
         which is the link back to experiment 5""")

    print("\n      why watch the node BELOW you, not the leader:")
    print(f"        {len(order) + 1} candidates all watching the leader means")
    print(f"        {len(order) + 1} clients woken by one failure -- the HERD EFFECT.")
    print("""        Watching your immediate predecessor wakes exactly ONE
        client per failure. The recipe is not arbitrary; it is a
        thundering-herd fix, and examiners like that you know why""")

    # Step 2: Take a distributed lock
    print("\n    a distributed lock is the SAME recipe:")
    zk.create("/locks")
    holders = []
    for client in ("jobA", "jobB"):
        p = zk.create("/locks/write-", data=client, ephemeral=True,
                      sequential=True, session=client)
        holders.append((client, p))
    first = zk.children("/locks")[0]
    print(f"      jobA and jobB both asked; {zk.tree[first]['data']} holds the lock")
    assert zk.tree[first]["data"] == "jobA"
    zk.expire("jobA")
    nxt = zk.children("/locks")[0]
    print(f"      jobA CRASHES -- lock passes to {zk.tree[nxt]['data']} automatically")
    assert zk.tree[nxt]["data"] == "jobB"
    print("""         the lock is released by the SESSION DYING, not by the
         client remembering to release it. A lock in a normal
         database survives the crash of whoever held it and deadlocks
         the system; an ephemeral znode cannot""")

    # Step 3: Rely on an atomic create
    print("\n    create is ATOMIC, which is the other half of every recipe:")
    zk.create("/config", data="v1")
    try:
        zk.create("/config", data="v2")
        raise AssertionError("the second create must fail")
    except FileExistsError as exc:
        print(f"      second create -> {type(exc).__name__}")
    print("""         exactly one client wins a create, cluster-wide, with no
         further negotiation. 'Whoever creates /master is the master'
         is a complete election algorithm in one line, and it works
         only because ZooKeeper linearises writes""")

    # Step 4: Size the ensemble
    print("\n    ensemble sizing -- why every cluster has an ODD number:")
    print(f"      {'servers':>8}{'quorum':>8}{'can lose':>10}  {'verdict'}")
    for n in (1, 2, 3, 4, 5, 6, 7):
        q = quorum(n)
        tol = n - q
        verdict = ("no fault tolerance" if tol == 0 else
                   "same tolerance as " + str(n - 1) if n % 2 == 0 else "good")
        print(f"      {n:>8}{q:>8}{tol:>10}  {verdict}")
    assert quorum(3) == 2 and quorum(4) == 3
    assert (3 - quorum(3)) == (4 - quorum(4)) == 1
    print("""         3 servers tolerate 1 failure. FOUR SERVERS ALSO TOLERATE
         ONE -- the extra machine buys nothing and adds a write to
         every quorum. That is the whole reason ZooKeeper ensembles
         are 3, 5 or 7, and it is a two-line exam answer""")

    # Step 5: See what ZooKeeper is not
    print("\n    what ZooKeeper is NOT:")
    print(f"      {'misuse':<34}{'why it fails'}")
    for m, w in (("a message queue", "no ordering guarantees across znodes"),
                 ("a data store", "1 MB per znode, whole tree in RAM"),
                 ("a cache", "every write is a quorum round trip"),
                 ("a service registry for 10k nodes", "watch storms")):
        print(f"      {m:<34}{w}")
    print("""         ZooKeeper stores COORDINATION STATE -- who is the leader,
         who holds the lock, what is the config -- and it is small,
         consistent and slow on purpose. Putting application data in
         it is the mistake that gets clusters into trouble""")

    # Step 6: Name who uses it
    print("\n    who uses it in this course:")
    for who, why in (("HDFS NameNode HA", "elects the ACTIVE NameNode (exp 5)"),
                     ("YARN ResourceManager HA", "elects the active RM (exp 6)"),
                     ("HBase", "tracks the master and the RegionServers (exp 15)"),
                     ("Kafka (pre-3.x)", "broker membership and controller")):
        print(f"      {who:<26}{why}")
    print("""         every HA story in the Hadoop ecosystem ends at
         ZooKeeper, which is why an experiment that looks like a
         detour is actually the keystone""")


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 16_zookeeper.sh:

OUTPUT

$ for id in 1 2 3; do
    mkdir -p zk$id && echo $id > zk$id/myid
    cat > zk$id.cfg <<EOF
  tickTime=2000
  initLimit=10
  syncLimit=5
  dataDir=$PWD/zk$id
  clientPort=218$id
  server.1=localhost:2888:3888
  server.2=localhost:2889:3889
  server.3=localhost:2890:3890
  4lw.commands.whitelist=srvr,mntr
  EOF
  done
$ for id in 1 2 3; do ZOO_LOG_DIR=zk$id zkServer.sh start $PWD/zk$id.cfg 2>&1 | tail -1; done
Starting zookeeper ... STARTED
Starting zookeeper ... STARTED
Starting zookeeper ... STARTED
$ sleep 10
$ for id in 1 2 3; do zkServer.sh status $PWD/zk$id.cfg 2>&1 | grep Mode; done
Mode: follower
Mode: leader
Mode: follower
$ echo srvr | nc localhost 2181 | grep -E "Mode|Zxid|Node count"
Zxid: 0x0
Mode: follower
Node count: 5
$ echo mntr | nc localhost 2181 | grep -E "zk_server_state|zk_znode_count|zk_outstanding_requests"
zk_server_state follower
zk_outstanding_requests 0
zk_znode_count  5
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
  ls /
  create /app "config-v1"
  get /app
  set /app "config-v2"
  stat /app
  ls -R /
  quit
  EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] ls /
[zookeeper]
[zk: localhost:2181(CONNECTED) 1] create /app "config-v1"
Created /app
[zk: localhost:2181(CONNECTED) 2] get /app
config-v1
[zk: localhost:2181(CONNECTED) 3] set /app "config-v2"
[zk: localhost:2181(CONNECTED) 4] stat /app
cZxid = 0x100000002
ctime = Sun Oct 04 23:31:34 UTC 2026
mZxid = 0x100000003
mtime = Sun Oct 04 23:31:34 UTC 2026
pZxid = 0x100000002
cversion = 0
dataVersion = 1
aclVersion = 0
ephemeralOwner = 0x0
dataLength = 9
numChildren = 0
[zk: localhost:2181(CONNECTED) 5] ls -R /
/
/app
/zookeeper
/zookeeper/config
/zookeeper/quota
[zk: localhost:2181(CONNECTED) 6] quit

WATCHER::

WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
  create -e /app/worker-1 "alive"
  ls /app
  quit
  EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create -e /app/worker-1 "alive"
Created /app/worker-1
[zk: localhost:2181(CONNECTED) 1] ls /app
[worker-1]
[zk: localhost:2181(CONNECTED) 2] quit

WATCHER::

WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2182 2>&1 <<'EOF'
  ls /app
  quit
  EOF
Connecting to localhost:2182
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2182(CONNECTED) 0] ls /app
[]
[zk: localhost:2182(CONNECTED) 1] quit

WATCHER::

WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
  create -s /app/task- "t"
  create -s /app/task- "t"
  create -s /app/task- "t"
  quit
  EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create -s /app/task- "t"
Created /app/task-0000000001
[zk: localhost:2181(CONNECTED) 1] create -s /app/task- "t"
Created /app/task-0000000002
[zk: localhost:2181(CONNECTED) 2] create -s /app/task- "t"
Created /app/task-0000000003
[zk: localhost:2181(CONNECTED) 3] quit

WATCHER::

WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
  create /election ""
  create -e -s /election/n- "nn1"
  create -e -s /election/n- "nn2"
  ls /election
  get -w /election/n-0000000000
  quit
  EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create /election ""
Created /election
[zk: localhost:2181(CONNECTED) 1] create -e -s /election/n- "nn1"
Created /election/n-0000000000
[zk: localhost:2181(CONNECTED) 2] create -e -s /election/n- "nn2"
Created /election/n-0000000001
[zk: localhost:2181(CONNECTED) 3] ls /election
[n-0000000000, n-0000000001]
[zk: localhost:2181(CONNECTED) 4] get -w /election/n-0000000000
nn1
[zk: localhost:2181(CONNECTED) 5] quit

WATCHER::

WatchedEvent state:SyncConnected type:NodeDeleted path:/election/n-0000000000

WATCHER::

WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
  ls /hbase
  ls /hadoop-ha/mycluster
  ls /rmstore
  quit
  EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled

WATCHER::

WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] ls /hbase
Node does not exist: /hbase
[zk: localhost:2181(CONNECTED) 1] ls /hadoop-ha/mycluster
Node does not exist: /hadoop-ha/mycluster
[zk: localhost:2181(CONNECTED) 2] ls /rmstore
Node does not exist: /rmstore
[zk: localhost:2181(CONNECTED) 3] quit

WATCHER::

WatchedEvent state:Closed type:None path:null
Exiting JVM with code 1
$ zkCli.sh -server localhost:2181 deleteall /app 2>/dev/null | tail -1
WatchedEvent state:SyncConnected type:None path:null
$ for id in 1 2 3; do zkServer.sh stop $PWD/zk$id.cfg 2>&1 | tail -1; done
Stopping zookeeper ... STOPPED
Stopping zookeeper ... STOPPED
Stopping zookeeper ... STOPPED

The Python check, 16_zookeeper_model.py:

OUTPUT

  Experiment 16 -- ZooKeeper coordination

    leader election, the standard recipe:
      every candidate creates an EPHEMERAL SEQUENTIAL znode
      the LOWEST sequence number is the leader
      everyone else watches the node just below them
      nn1 -> lock-0000000000
      nn2 -> lock-0000000001
      nn3 -> lock-0000000002

      LEADER: nn1

      nn1's session expires (its JVM was killed):
      lock-0000000000 vanished; new LEADER: nn2
         nobody ran a failover script. The ephemeral node was
         deleted BY THE SERVER when the heartbeat stopped, the watch
         fired, and nn2 saw itself at the head of the queue. That is
         how HDFS NameNode HA actually chooses its active node --
         which is the link back to experiment 5

      why watch the node BELOW you, not the leader:
        3 candidates all watching the leader means
        3 clients woken by one failure -- the HERD EFFECT.
        Watching your immediate predecessor wakes exactly ONE
        client per failure. The recipe is not arbitrary; it is a
        thundering-herd fix, and examiners like that you know why

    a distributed lock is the SAME recipe:
      jobA and jobB both asked; jobA holds the lock
      jobA CRASHES -- lock passes to jobB automatically
         the lock is released by the SESSION DYING, not by the
         client remembering to release it. A lock in a normal
         database survives the crash of whoever held it and deadlocks
         the system; an ephemeral znode cannot

    create is ATOMIC, which is the other half of every recipe:
      second create -> FileExistsError
         exactly one client wins a create, cluster-wide, with no
         further negotiation. 'Whoever creates /master is the master'
         is a complete election algorithm in one line, and it works
         only because ZooKeeper linearises writes

    ensemble sizing -- why every cluster has an ODD number:
       servers  quorum  can lose  verdict
             1       1         0  no fault tolerance
             2       2         0  no fault tolerance
             3       2         1  good
             4       3         1  same tolerance as 3
             5       3         2  good
             6       4         2  same tolerance as 5
             7       4         3  good
         3 servers tolerate 1 failure. FOUR SERVERS ALSO TOLERATE
         ONE -- the extra machine buys nothing and adds a write to
         every quorum. That is the whole reason ZooKeeper ensembles
         are 3, 5 or 7, and it is a two-line exam answer

    what ZooKeeper is NOT:
      misuse                            why it fails
      a message queue                   no ordering guarantees across znodes
      a data store                      1 MB per znode, whole tree in RAM
      a cache                           every write is a quorum round trip
      a service registry for 10k nodes  watch storms
         ZooKeeper stores COORDINATION STATE -- who is the leader,
         who holds the lock, what is the config -- and it is small,
         consistent and slow on purpose. Putting application data in
         it is the mistake that gets clusters into trouble

    who uses it in this course:
      HDFS NameNode HA          elects the ACTIVE NameNode (exp 5)
      YARN ResourceManager HA   elects the active RM (exp 6)
      HBase                     tracks the master and the RegionServers (exp 15)
      Kafka (pre-3.x)           broker membership and controller
         every HA story in the Hadoop ecosystem ends at
         ZooKeeper, which is why an experiment that looks like a
         detour is actually the keystone

The script starts and stops its own ensemble, three servers on one machine. Which of them wins the election differs from run to run; exactly one says Mode: leader.

WATCH YOUR PREDECESSOR, NOT THE LEADER

Watching the leader wakes every candidate on one failure — the herd effect. Watching your immediate predecessor wakes exactly one.

The lock is the same recipe.

jobA holds the lock
jobA CRASHES -- lock passes to jobB automatically

Released by the session dying, not by the client remembering. A database lock survives its holder's crash and deadlocks the system; an ephemeral znode cannot.

ENSEMBLE SIZING

Servers Quorum Can lose Verdict
3 2 1 good
4 3 1 same as 3
5 3 2 good
6 4 2 same as 5
7 4 3 good

Four servers tolerate one failure — exactly what three tolerate. The fourth machine buys nothing and adds a write to every quorum. 3, 5 or 7.

RESULT

Three servers elected one leader; every server held the same tree; an ephemeral node vanished when its session ended; the sequence numbers started at 1, not 0, because the parent counts every child it has had. In the model, four servers tolerate one failure, as three do.

Experiment 17 — Spark over HBase

1. Question

Process HBase datasets using Spark integration with Hadoop.

2. Aim

Read an HBase table into Spark as an RDD, a partition per region, and query it with Spark SQL; then, in PySpark, count words with RDDs, compare reduceByKey with groupByKey, query the star schema and cache.

3. Steps

On the cluster, 17_spark_hbase.scala:

  1. Configure the HBase connection and scan.
  2. Read the table as an RDD, a partition per region.
  3. Map each row to a case class.
  4. Query it with Spark SQL.

The Python check, 17_spark.py:

  1. Start a SparkSession.
  2. Count words with RDDs.
  3. Compare reduceByKey with groupByKey.
  4. See lazy evaluation and the DAG.
  5. Query the star schema with DataFrames.
  6. Analyse the logs.
  7. Compare Spark with MapReduce.
  8. Cache, and measure it.

TWO REAL ENGINES

real SparkSession: version 4.2.0, master local[2]

PySpark installs from PyPI and Java 21 is present, so a genuine session starts, real RDDs are built, and a real shuffle happens inside reduceByKey. The Scala half runs in spark-shell against HBase, with HBase's own jars on the classpath and its TableInputFormat — the connector Hadoop ships — so no separate connector is needed.

4. Programme

On the cluster, 17_spark_hbase.scala:

// Experiment 17 -- process HBase datasets using Spark integration with Hadoop
//
// Run it: the spark-shell command below, with HBase running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 17_spark.py, which runs REAL PySpark -- only the HBase connector is missing
//
// run with:
//   spark-shell --master 'local[*]' \
//     --jars $(ls $HBASE_HOME/lib/*.jar | tr '\n' ',') \
//     -i 17_spark_hbase.scala
// [Corrected: the jars were in /usr/lib/hbase/lib, where some distributions put
// HBase; $HBASE_HOME is wherever yours is. `hbase mapredcp` lists the smaller
// set a job needs, and works the same. And --master yarn needs YARN containers
// running Java 17 for Spark 4; this cluster's run Java 8, so this runs
// local[*] -- the same code, on one machine.]

import org.apache.hadoop.hbase.{HBaseConfiguration, CellUtil}
import org.apache.hadoop.hbase.client.Result
import org.apache.hadoop.hbase.io.ImmutableBytesWritable
import org.apache.hadoop.hbase.mapreduce.TableInputFormat
import org.apache.hadoop.hbase.util.Bytes

// Step 1: Configure the HBase connection and scan
val conf = HBaseConfiguration.create()
conf.set("hbase.zookeeper.quorum", "localhost")        // ZooKeeper, again
// [Corrected: the quorum was zk1,zk2,zk3, experiment 16's three hosts. HBase
// here runs its own ZooKeeper, on localhost.]
conf.set(TableInputFormat.INPUT_TABLE, "sales")
// push the scan down: read one region, not the table
conf.set(TableInputFormat.SCAN_ROW_START, "South")
conf.set(TableInputFormat.SCAN_ROW_STOP,  "South~")
conf.set(TableInputFormat.SCAN_COLUMNS,   "sales:revenue sales:qty")

// Step 2: Read the table as an RDD, a partition per region
val hBaseRDD = sc.newAPIHadoopRDD(
  conf,
  classOf[TableInputFormat],
  classOf[ImmutableBytesWritable],
  classOf[Result])

// ONE SPARK PARTITION PER HBASE REGION. That is the whole integration:
// Spark reads regions in parallel, locally, without going through the
// RegionServer's RPC path for bulk scans.
println(s"partitions = ${hBaseRDD.getNumPartitions}")
// [Noted, from running it: straight after the puts this said partitions = 0,
// and every query below came back empty, with no error. The rows were still
// in the memstore, in no store file, so the region's size was 0 -- and
// TableInputFormat made no split for it. `flush 'sales'` in the HBase shell
// first; on a busy table the memstore flushes by itself.]

// Step 3: Map each row to a case class
case class Sale(rowKey: String, region: String, qty: Int, revenue: Double)

val sales = hBaseRDD.map { case (_, result) =>
  val key = Bytes.toString(result.getRow)
  val qty = Option(result.getValue(Bytes.toBytes("sales"), Bytes.toBytes("qty")))
              .map(b => Bytes.toString(b).toInt).getOrElse(0)
  val rev = Option(result.getValue(Bytes.toBytes("sales"), Bytes.toBytes("revenue")))
              .map(b => Bytes.toString(b).toDouble).getOrElse(0.0)
  Sale(key, key.split("#")(0), qty, rev)
}

// Step 4: Query it with Spark SQL
import spark.implicits._
val df = sales.toDF()
df.createOrReplaceTempView("sales")

spark.sql("""
  SELECT region, SUM(revenue) AS revenue, SUM(qty) AS units
  FROM sales GROUP BY region ORDER BY revenue DESC
""").show()
//   expected: South 10360 -- the same number Course 11's DAX, experiment 10's
//   SQL and experiment 17's PySpark all produce.
// [Corrected: this expected North 2520 as well. The scan above reads only the
// rows from 'South' to 'South~' -- that is the point of pushing it down -- so
// North is never read, and the query has one row.]

// --- writing BACK to HBase, in bulk ---------------------------------------
// Never use put() per row from a Spark job: that is one RPC per record and it
// will overwhelm the RegionServers. Write HFiles and load them:
//
//   df.rdd.map(toKeyValue).sortByKey()
//     .saveAsNewAPIHadoopFile(path, classOf[ImmutableBytesWritable],
//                             classOf[KeyValue], classOf[HFileOutputFormat2], conf)
//   LoadIncrementalHFiles.doBulkLoad(new Path(path), admin, table, locator)
//
// Bulk load bypasses the write path entirely -- no WAL, no memstore, no
// flush -- and is one to two orders of magnitude faster than put().

// --- when NOT to do this ---------------------------------------------------
// A full-table Spark scan of HBase is SLOWER than the same data in Parquet,
// because HBase stores every cell with its row key, family, qualifier and
// timestamp. HBase is for random reads and writes; Parquet is for scans.
// If every job you run is a full scan, the data is in the wrong store.

The Python check, 17_spark.py:

"""Experiment 17 -- process HBase datasets using Spark integration with Hadoop.

THIS EXPERIMENT RUNS REAL SPARK. PySpark 4.2 installs from PyPI and Java 21 is
present, so a genuine SparkSession starts, real RDDs are built, and a real
shuffle happens inside reduceByKey. Nothing here is a simulation.

The dataset here comes from the experiment 15 model rather than an HBase table;
`17_spark_hbase.scala` reads a real one, through TableInputFormat, in
spark-shell (the lab page shows it).

Run with:  /tmp/sparkenv/bin/python 17_spark.py
or let tools/run_bigdata_labs.py find the environment for you.
"""
import os
import sys

import fixtures as f


def spark_available():
    try:
        import pyspark  # noqa: F401
        return True
    except ImportError:
        return False


def main():
    print("  Experiment 17 -- Spark on the Hadoop stack")

    if not spark_available():
        print("""
    PySpark is not importable from this interpreter.
    Run tools/setup_spark.sh, then use /tmp/sparkenv/bin/python.
    SKIPPED -- and this line is what a skipped experiment looks like.""")
        return False

    os.environ.setdefault("PYSPARK_PYTHON", sys.executable)
    from pyspark.sql import SparkSession
    from pyspark.sql import functions as F

    # Step 1: Start a SparkSession
    spark = (SparkSession.builder
             .appName("course-12b-exp-17")
             .master("local[2]")
             .config("spark.ui.enabled", "false")
             .config("spark.sql.shuffle.partitions", "4")
             .getOrCreate())
    spark.sparkContext.setLogLevel("ERROR")
    print(f"\n    real SparkSession: version {spark.version}, "
          f"master {spark.sparkContext.master}")

    # Step 2: Count words with RDDs
    rdd = spark.sparkContext.parallelize(list(f.DOCS.values()), 3)
    counts = (rdd.flatMap(lambda line: line.split())
                 .map(lambda w: (w, 1))
                 .reduceByKey(lambda a, b: a + b))
    got = dict(counts.collect())
    print(f"\n    RDD word count: {len(got)} distinct words, "
          f"{sum(got.values())} total")
    assert sum(got.values()) == 48
    assert got["the"] == 5 and got["big"] == 4 and got["dog"] == 4
    print(f"      partitions: input {rdd.getNumPartitions()}, "
          f"after reduceByKey {counts.getNumPartitions()}")
    print("""         IDENTICAL to experiment 7's MapReduce answer, on a real
         distributed engine. reduceByKey is map -> COMBINE -> shuffle
         -> reduce; Spark applies the combiner automatically, which
         MapReduce makes you ask for""")

    # Step 3: Compare reduceByKey with groupByKey
    print("\n    reduceByKey against groupByKey -- the same answer, not the same job:")
    grouped = (rdd.flatMap(lambda line: line.split())
                  .map(lambda w: (w, 1))
                  .groupByKey()
                  .mapValues(len))
    assert dict(grouped.collect()) == got
    # measure the map-side combine for THIS partitioning rather than quoting
    # experiment 7's number, which was per-document and not per-partition
    per_part = (rdd.flatMap(lambda line: line.split())
                   .map(lambda w: (w, 1))
                   .mapPartitions(lambda it: [len({k for k, _ in it})])
                   .collect())
    combined = sum(per_part)
    print(f"      groupByKey  : shuffles all 48 pairs, then counts")
    print(f"      reduceByKey : combines to {combined} map-side "
          f"({per_part} per partition), then shuffles")
    assert combined < 48
    print("""         same output, and groupByKey moves every record across
         the network while reduceByKey moves one per key per
         partition. On a real corpus groupByKey is how you produce an
         OutOfMemoryError on a single hot key. This is the most
         examined Spark question there is""")

    # Step 4: See lazy evaluation and the DAG
    lineage = (rdd.flatMap(lambda l: l.split())
                  .filter(lambda w: len(w) > 3)
                  .map(lambda w: (w[0], 1)))
    print(f"\n    lazy evaluation: three transformations queued, nothing ran")
    print(f"      the DAG has {len(lineage.toDebugString().decode().splitlines())} "
          f"stages of lineage recorded")
    result = lineage.reduceByKey(lambda a, b: a + b).collect()
    print(f"      .collect() is the ACTION -- it returned {len(result)} keys")
    assert len(result) > 0
    print("""         transformations build a DAG; only an ACTION submits it.
         That is why a typo in a map() surfaces at collect() and not
         where you wrote it -- and why Spark can fuse the whole chain
         into one pass over the data""")

    # Step 5: Query the star schema with DataFrames
    sdf = spark.createDataFrame(f.SALES_DF)
    agg = (sdf.groupBy("region")
              .agg(F.sum("revenue").alias("revenue"),
                   F.sum("profit").alias("profit"))
              .orderBy(F.desc("revenue")))
    rows = {r["region"]: r["revenue"] for r in agg.collect()}
    print(f"\n    DataFrame aggregate over the SAME nine rows:")
    for r in agg.collect():
        print(f"      {r['region']:<8}{r['revenue']:>10,.0f}{r['profit']:>9,.0f}")
    assert rows["South"] == 10360.0 and rows["North"] == 2520.0
    assert sum(rows.values()) == f.total_revenue()
    print("""         10,360 and 2,520 again -- the third engine to produce
         them, after Course 11's DAX and experiment 10's SQL. Spark,
         DuckDB and Power BI agree, which is what reusing one dataset
         across three courses was for""")

    # Step 6: Analyse the logs
    logs = spark.sparkContext.parallelize(f.access_logs(40), 2)
    by_status = (logs.map(lambda ln: (ln.rsplit(" ", 2)[-2], 1))
                     .reduceByKey(lambda a, b: a + b)
                     .collectAsMap())
    print(f"\n    the ingested access logs, aggregated in Spark:")
    for code in sorted(by_status):
        print(f"      HTTP {code}: {by_status[code]}")
    assert by_status["200"] == 24 and by_status["404"] == 8
    assert sum(by_status.values()) == 40
    print("""         the same 24 / 8 / 8 the Flume agent produced in
         experiment 12. Ingest with Flume, analyse with Spark, on
         bytes that were never transformed in between -- that is the
         end-to-end story the syllabus asks for""")

    # Step 7: Compare Spark with MapReduce
    print("\n    Spark against MapReduce, on the parts that decided it:")
    print(f"      {'':<22}{'MapReduce':<26}{'Spark'}")
    for label, mr, sp in (
            ("between stages", "writes to HDFS", "keeps in MEMORY"),
            ("iterative jobs", "re-reads every pass", "cache() once"),
            ("API", "map and reduce only", "~80 operators"),
            ("interactive", "no", "yes -- the shell"),
            ("fault tolerance", "re-run the task", "recompute from LINEAGE"),
            ("streaming", "no", "structured streaming"),
            ("runs on YARN", "yes", "yes -- same cluster")):
        print(f"      {label:<22}{mr:<26}{sp}")
    print("""         the decisive row is the first. A ten-iteration machine
         learning job writes to HDFS nine times under MapReduce and
         zero times under Spark, which is where the '100x faster'
         headline comes from -- it is a claim about ITERATIVE jobs,
         and quoting it for a single-pass job is wrong""")

    # Step 8: Cache, and measure it
    base = spark.sparkContext.parallelize(range(200_000), 4).map(lambda x: x * 2)
    base.cache()
    first = base.sum()
    second = base.sum()
    assert first == second == sum(x * 2 for x in range(200_000))
    print(f"\n    cache(): two actions over the same RDD, sum = {first:,}")
    print(f"      storage level after cache(): {base.getStorageLevel()}")
    print("""         WITHOUT cache() the second sum recomputes the map from
         the source. With it, only the first action pays. Caching is
         the single highest-value Spark optimisation and the one
         students forget, because nothing FAILS without it -- the job
         is merely twice as slow""")
    base.unpersist()

    spark.stop()
    print("\n    SparkSession stopped cleanly.")
    return True


if __name__ == "__main__":
    main()

5. Execution and Results

On the cluster, 17_spark_hbase.scala:

OUTPUT

$ start-hbase.sh > /dev/null 2>&1
$ sleep 20
$ hbase shell -n < load_sales.hbase 2>&1 | grep -E 'row\(s\)$|^=> [0-9]+$'
9 row(s)
=> 9
$ JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64 $SPARK_HOME/bin/spark-shell --master 'local[*]' --conf spark.ui.enabled=false --jars $(ls $HBASE_HOME/lib/*.jar | tr '\n' ',') -i 17_spark_hbase.scala < /dev/null 2>spark.log
Welcome to
      ____              __
     / __/__  ___ _____/ /__
    _\ \/ _ \/ _ `/ __/  '_/
   /___/ .__/\_,_/_/ /_/\_\   version 4.2.0
      /_/

Using Scala version 2.13.18 (OpenJDK 64-Bit Server VM, Java 21.0.10)
Type in expressions to have them evaluated.
Type :help for more information.
Spark context available as 'sc' (master = local[*], app id = local-1791156759020).
Spark session available as 'spark'.
partitions = 1
+------+-------+-----+
|region|revenue|units|
+------+-------+-----+
| South|10360.0|   48|
+------+-------+-----+


scala> :quit
$ awk '/ERROR|Exception/' spark.log | sed -E 's/^[0-9/]+ [0-9:]+ //' | sort -u | head -5
$ stop-hbase.sh 2>&1 | tail -1
stopping hbase.............

The Python check, 17_spark.py:

OUTPUT

  Experiment 17 -- Spark on the Hadoop stack

    real SparkSession: version 4.2.0, master local[2]

    RDD word count: 26 distinct words, 48 total
      partitions: input 3, after reduceByKey 3
         IDENTICAL to experiment 7's MapReduce answer, on a real
         distributed engine. reduceByKey is map -> COMBINE -> shuffle
         -> reduce; Spark applies the combiner automatically, which
         MapReduce makes you ask for

    reduceByKey against groupByKey -- the same answer, not the same job:
      groupByKey  : shuffles all 48 pairs, then counts
      reduceByKey : combines to 35 map-side ([11, 14, 10] per partition), then shuffles
         same output, and groupByKey moves every record across
         the network while reduceByKey moves one per key per
         partition. On a real corpus groupByKey is how you produce an
         OutOfMemoryError on a single hot key. This is the most
         examined Spark question there is

    lazy evaluation: three transformations queued, nothing ran
      the DAG has 2 stages of lineage recorded
      .collect() is the ACTION -- it returned 11 keys
         transformations build a DAG; only an ACTION submits it.
         That is why a typo in a map() surfaces at collect() and not
         where you wrote it -- and why Spark can fuse the whole chain
         into one pass over the data

    DataFrame aggregate over the SAME nine rows:
      South       10,360    2,760
      North        2,520      765
         10,360 and 2,520 again -- the third engine to produce
         them, after Course 11's DAX and experiment 10's SQL. Spark,
         DuckDB and Power BI agree, which is what reusing one dataset
         across three courses was for

    the ingested access logs, aggregated in Spark:
      HTTP 200: 24
      HTTP 404: 8
      HTTP 500: 8
         the same 24 / 8 / 8 the Flume agent produced in
         experiment 12. Ingest with Flume, analyse with Spark, on
         bytes that were never transformed in between -- that is the
         end-to-end story the syllabus asks for

    Spark against MapReduce, on the parts that decided it:
                            MapReduce                 Spark
      between stages        writes to HDFS            keeps in MEMORY
      iterative jobs        re-reads every pass       cache() once
      API                   map and reduce only       ~80 operators
      interactive           no                        yes -- the shell
      fault tolerance       re-run the task           recompute from LINEAGE
      streaming             no                        structured streaming
      runs on YARN          yes                       yes -- same cluster
         the decisive row is the first. A ten-iteration machine
         learning job writes to HDFS nine times under MapReduce and
         zero times under Spark, which is where the '100x faster'
         headline comes from -- it is a claim about ITERATIVE jobs,
         and quoting it for a single-pass job is wrong

    cache(): two actions over the same RDD, sum = 39,999,800,000
      storage level after cache(): Memory Serialized 1x Replicated
         WITHOUT cache() the second sum recomputes the map from
         the source. With it, only the first action pays. Caching is
         the single highest-value Spark optimisation and the one
         students forget, because nothing FAILS without it -- the job
         is merely twice as slow

    SparkSession stopped cleanly.

_drive_17_spark_hbase.py starts HBase as for experiment 15, puts the nine sales rows into sales with the row key region#store#date#product, and flushes the table — a scan through TableInputFormat reads the store files, and the rows still in memory were not seen until the flush; the file says so. spark-shell runs on Java 21 and HBase on Java 8.

17_spark.py's output is what it printed to stdout. Spark's JVM logs to stderr — a timestamped line per event and a progress bar — and that is left out. Among it is one warning worth knowing: PySpark 4.2 "does not yet fully support pandas >= 3.0.0". The program uses pandas only to load the shared fixtures, and every figure it prints is asserted.

The RDD word count. 26 distinct words, 48 total — identical to experiment 7's MapReduce answer, on a real distributed engine.

REDUCEBYKEY AGAINST GROUPBYKEY

What crosses the network
groupByKey all 48 pairs, then counts
reduceByKey combines to 35 map-side — [11, 14, 10] per partition

WHY IT MATTERS

Note 35, not experiment 7's 39. Spark's three partitions each hold two documents, so more merging happens per task. The combiner's saving depends on the split, exactly as Unit 3 said — and this is the same measurement made two ways.

Lazy evaluation. Three transformations queued; .collect() is the action that submits the DAG. That is why a typo in a map() surfaces at collect().

THE CROSS-COURSE CHECK, FOURTH ENGINE

Region Revenue Profit
South 10,360 2,760
North 2,520 765

Business Intelligence Tools' DAX, DuckDB, Hive and Spark all produce ₹10,360.

The logs from experiment 12. HTTP 200: 24, 404: 8, 500: 8 — the same 24 / 8 / 8 the Flume agent produced. Ingest with Flume, analyse with Spark, on bytes never transformed in between.

cache().

two actions over the same RDD, sum = 39,999,800,000
storage level: Memory Serialized 1x Replicated

Without cache() the second action recomputes the whole lineage. Caching is the highest-value Spark optimisation and the one students forget — because nothing fails without it; the job is merely twice as slow.

THE SPARK-OVER-HBASE CAVEAT

One Spark partition per HBase region is the whole integration — the nine rows are one region, so one partition. But never put() per row from a Spark job — write HFiles and bulk-load them.

And: a full-table Spark scan of HBase is slower than the same data in Parquet, because HBase stores every cell with its row key, family, qualifier and timestamp. If every job is a full scan, the data is in the wrong store.

RESULT

Spark read the nine rows from HBase in one partition — one region — and Spark SQL gave South ₹10,360 from 48 units, as DAX, Hive and DuckDB did. The scan saw nothing until the table was flushed. In PySpark the word count is 26 words and 48 in all, and reduceByKey shuffles 35 records against groupByKey's 48.


What the runner asserts

Script Experiments Real tool?
04_blocks_replication.py 4 arithmetic
05_fault_tolerance.py 5 model
06_yarn_scheduling.py 6 model
07_wordcount.py 7 explicit MapReduce engine
08_inverted_index.py 8 same engine
09_pig_equivalent.py 9 dataflow, step by step
10_hive_duckdb.py 10 real SQL, DuckDB
11_sqoop_equivalent.py 11 real SQLite + real Parquet
12_flume_equivalent.py 12 agent semantics
13_avro_parquet.py 13 real Avro + real Parquet
14_pipeline.py 14 real end-to-end
15_hbase_model.py 15 model
16_zookeeper_model.py 16 model
17_spark.py 17 REAL APACHE SPARK

Plus the 15 tool files, run on the cluster, each checked for the answers it must print — Pig's three categories, Hive's South ₹10,360, Sqoop's 90 imported and 3 exported, Flume's 8 routed events — and none may still say NOT EXECUTED.

Experiment 17's PySpark half skips loudly if the PySpark environment is absent, and the tool files are only audited if the Hadoop stack is — the same graceful-skip pattern Web Technologies uses for jsdom. A skip is not a pass, and the runner says so.


Lab examination

Two hours on a cluster, one experiment number, then a viva.

What costs marks:

What earns them:

Each program, on its own page

The same experiments, one page each, so a program can be reached by what it does rather than by its number.

RUNS

Store and retrieve a large file in HDFS: blocks

RUNS

Simulate NameNode/DataNode failure and observe fault in Hadoop and Spark

RUNS

Configure YARN, run sample applications

RUNS

A simple MapReduce program for word count

RUNS

A MapReduce job for inverted index creation

RUNS

Data analysis with Pig Latin scripts in Hadoop and Spark

RUNS

Hive queries for structured data analysis: tables

RUNS

Import data from an RDBMS into Hadoop using Sqoop

RUNS

Capture and store log/streaming data using Flume in Hadoop and Spark

RUNS

Serialize and store datasets in Avro and Parquet in Hadoop and Spark

RUNS

An end-to-end ingestion workflow combining batch (Sqoop) in Hadoop and Spark

RUNS

Create and manage tables in HBase (CRUD operations) in Hadoop and Spark

RUNS

Demonstrate coordination with ZooKeeper in Hadoop and Spark

RUNS

Process HBase datasets using Spark integration with Hadoop