17 experiments, each set out as 1. Question, 2. Aim, 3. Steps, 4. Programme, 5. Execution and Results.
Code lives in labs/course-12b-bigdata/.
Every file in this lab runs, on a real Hadoop 3.3.6 cluster. Most experiments have two halves:
| Half | Files | Status |
|---|---|---|
| The tool you run on a cluster | 15 files — shell scripts of hdfs, yarn, sqoop and zkCli.sh commands, Pig, HiveQL, the HBase shell, a Flume agent, Java and Scala |
Run, each on a fresh cluster, by tools/data-science/hadoop_lab.py; what it printed is under 5. Execution and Results |
| The check | 14 programs | Executed and asserted by tools/data-science/run_bigdata_labs.py, which also runs the tool files and checks their answers |
(Updated October 2026: Hadoop, Pig, Hive, HBase, ZooKeeper, Sqoop and Flume could not be installed where these labs are checked, and every tool file said NOT EXECUTED. They now install from archive.apache.org, with MariaDB from the Ubuntu archive for Sqoop, and running the files found faults that reading them had not — a Pig keyword used as a field name, two Hive statements that stop the script, a Sqoop query with two columns of one name, a Flume agent that dropped its HDFS sink and routed nothing, an HBase filter that let rows through, a Spark scan that saw no rows. Each is corrected in its file and noted under its experiment.)
bash tools/data-science/setup_hadoop.sh # Hadoop, Pig, Hive, HBase, ZooKeeper, Sqoop, Flume
bash tools/data-science/setup_spark.sh # PySpark, for experiment 17
pip install -r tools/requirements.txt
python3 tools/data-science/run_bigdata_labs.py # about half an hour; --audit-only skips the cluster
HOW THE TOOL FILES WERE RUN
hadoop_lab.py formats and starts a cluster in a temporary folder for each file — a NameNode,
four DataNodes (so replication 3 can survive the loss of one, and a block can be
re-replicated), a SecondaryNameNode, a ResourceManager, a NodeManager and the JobHistory server
— then runs the file as a student types it, printing $ command before each command, and stops
everything at the end.
Where a file assumes something is already there — the sales file in HDFS, a database for Sqoop,
HBase running — a _drive_ script beside it does that first, and those commands are printed
too, so the transcript shows everything that ran. One setting differs from a default cluster,
and the experiment that depends on it says so: a DataNode is declared dead after 60 s, not
630 s.
A cluster makes new ids, ports and times on every run — block ids, application ids, the port
a DataNode took, how long a job ran. The output shown is one run's, unchanged;
capture_lab_outputs.py --check compares a rerun with those values masked, and every other
line must repeat exactly.
Experiments 10, 14 and 17 all use Business Intelligence Tools' star schema, imported rather
than copied. fixtures.py loads labs/course-11-bi/fixtures.py by path at
import time, so the two courses cannot drift.
South = ₹10,360 is produced by Business Intelligence Tools' DAX CALCULATE, by DuckDB and by
Hive in experiment 10, and by Spark — reading HBase — in experiment 17. If they ever
disagree, one of them is wrong and verify_all.sh says so.
Install Hadoop on one machine, configure it, and start its daemons.
Download Hadoop, write its four configuration files, format HDFS, start the five daemons and check each one is up.
On the cluster, 01_install_hadoop.sh:
THE THREE INSTALLATION FAILURES EVERYONE HITS
JAVA_HOME not set inside hadoop-env.sh. Exporting it in your shell
is not enough — the daemons do not inherit it.
Re-running hdfs namenode -format after storing data. The DataNode's
clusterID no longer matches the NameNode's and it refuses to start. Fix:
delete the datanode directory, or edit its VERSION file.
ssh localhost prompting for a password. start-dfs.sh hangs for ever.
jps should show five processes: NameNode, DataNode, SecondaryNameNode,
ResourceManager, NodeManager.
On the cluster, 01_install_hadoop.sh:
# Experiment 1 -- installation and setup of a Hadoop single-node cluster
#
# Run it: bash 01_install_hadoop.sh, on Linux with Java 11 installed. It was run where these
# labs are checked, in an empty home directory, and the lab page shows what it printed.
# [Changed: this said the file had never been run, as Hadoop could not be installed there. It
# installs from archive.apache.org, and the corrections below are what running it found.]
#
# The runnable half is none -- installation has no query logic to verify
#
# Step 1: Check the prerequisites
java -version # Hadoop 3.x needs Java 8 or 11
# On your own machine, run Hadoop as its own user, with passwordless ssh to
# localhost -- start-dfs.sh and start-yarn.sh ssh to every host they start a
# daemon on, even in "single node":
# sudo adduser hadoop && su - hadoop
# ssh-keygen -t rsa -P '' -f ~/.ssh/id_rsa
# cat ~/.ssh/id_rsa.pub >> ~/.ssh/authorized_keys
# chmod 600 ~/.ssh/authorized_keys
# ssh localhost # must succeed WITHOUT a password
# Where this was run there is no ssh server, so the daemons are started one by
# one below, with the same commands start-dfs.sh runs on each host.
# Step 2: Download and unpack Hadoop
[ -f hadoop-3.3.6.tar.gz ] || \
curl -fO https://archive.apache.org/dist/hadoop/common/hadoop-3.3.6/hadoop-3.3.6.tar.gz
tar -xzf hadoop-3.3.6.tar.gz && mv hadoop-3.3.6 ~/hadoop
# [Corrected: the download was from dlcdn.apache.org, which keeps only the
# current releases -- 3.3.6 is no longer there, and wget got a 404. Every
# release stays in archive.apache.org. And it was moved to /usr/local/hadoop
# with sudo; your home directory needs no sudo.]
cat >> ~/.bashrc <<'EOF'
export HADOOP_HOME=~/hadoop
export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64
export PATH=$PATH:$HADOOP_HOME/bin:$HADOOP_HOME/sbin
export HADOOP_CONF_DIR=$HADOOP_HOME/etc/hadoop
EOF
source ~/.bashrc
# Step 3: Write the four configuration files
cat > $HADOOP_CONF_DIR/core-site.xml <<'EOF'
<configuration>
<property><name>fs.defaultFS</name><value>hdfs://localhost:9000</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/hdfs-site.xml <<EOF
<configuration>
<!-- 1, not 3: there is only one node -->
<property><name>dfs.replication</name><value>1</value></property>
<property><name>dfs.namenode.name.dir</name><value>file://$HOME/hadoop_store/hdfs/namenode</value></property>
<property><name>dfs.datanode.data.dir</name><value>file://$HOME/hadoop_store/hdfs/datanode</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/mapred-site.xml <<EOF
<configuration>
<property><name>mapreduce.framework.name</name><value>yarn</value></property>
<property><name>yarn.app.mapreduce.am.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
<property><name>mapreduce.map.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
<property><name>mapreduce.reduce.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
</configuration>
EOF
cat > $HADOOP_CONF_DIR/yarn-site.xml <<'EOF'
<configuration>
<property><name>yarn.nodemanager.aux-services</name><value>mapreduce_shuffle</value></property>
</configuration>
EOF
echo "export JAVA_HOME=$JAVA_HOME" >> $HADOOP_CONF_DIR/hadoop-env.sh
# [Corrected: the four files were listed as property names only; they are
# written out here, under your home directory rather than /usr/local, which
# needs no sudo. Hadoop 3 also needs HADOOP_MAPRED_HOME passed to MapReduce's
# containers, or every job fails with "Could not find or load main class
# org.apache.hadoop.mapreduce.v2.app.MRAppMaster".]
# Step 4: Format HDFS, and start the five daemons
hdfs namenode -format 2>&1 | grep "has been successfully formatted"
# ONCE. Re-formatting destroys the cluster. (It prints a hundred log lines;
# this is the one that says it worked.)
hdfs --daemon start namenode
hdfs --daemon start datanode
hdfs --daemon start secondarynamenode
yarn --daemon start resourcemanager
yarn --daemon start nodemanager
# = start-dfs.sh and start-yarn.sh, which run exactly these on each host,
# over ssh
sleep 15 # the DataNode and NodeManager register
jps | awk '{print $2}' | sort # expect: NameNode, DataNode,
# SecondaryNameNode, ResourceManager,
# NodeManager -- five processes, and Jps
hdfs dfsadmin -report 2>/dev/null | grep "Live datanodes"
hdfs dfs -mkdir -p /user/$USER && hdfs dfs -ls /user
# web UIs -- 200 means each is up:
curl -sL -o /dev/null -w "NameNode http://localhost:9870 %{http_code}\n" http://localhost:9870
curl -s -o /dev/null -w "ResourceManager http://localhost:8088 %{http_code}\n" http://localhost:8088/cluster
# Step 5: Stop them
yarn --daemon stop nodemanager
yarn --daemon stop resourcemanager
hdfs --daemon stop secondarynamenode
hdfs --daemon stop datanode
hdfs --daemon stop namenode
# = stop-yarn.sh and stop-dfs.sh
# --- the three failures everyone hits --------------------------------------
# 1. JAVA_HOME not set INSIDE hadoop-env.sh (the shell export is not enough)
# 2. re-running `hdfs namenode -format` after storing data: the DataNode's
# clusterID no longer matches the NameNode's, and the DataNode will not
# start. Fix: delete the datanode directory, or edit its VERSION file.
# 3. ssh localhost prompting for a password -- start-dfs.sh hangs for ever
On the cluster, 01_install_hadoop.sh:
OUTPUT
$ java -version
openjdk version "11.0.32.1" 2026-08-18
OpenJDK Runtime Environment (build 11.0.32.1+1-post-1ubuntu1-24.04-Ubuntu)
OpenJDK 64-Bit Server VM (build 11.0.32.1+1-post-1ubuntu1-24.04-Ubuntu, mixed mode, sharing)
$ [ -f hadoop-3.3.6.tar.gz ] || \
curl -fO https://archive.apache.org/dist/hadoop/common/hadoop-3.3.6/hadoop-3.3.6.tar.gz
$ tar -xzf hadoop-3.3.6.tar.gz && mv hadoop-3.3.6 ~/hadoop
$ cat >> ~/.bashrc <<'EOF'
export HADOOP_HOME=~/hadoop
export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64
export PATH=$PATH:$HADOOP_HOME/bin:$HADOOP_HOME/sbin
export HADOOP_CONF_DIR=$HADOOP_HOME/etc/hadoop
EOF
$ source ~/.bashrc
$ cat > $HADOOP_CONF_DIR/core-site.xml <<'EOF'
<configuration>
<property><name>fs.defaultFS</name><value>hdfs://localhost:9000</value></property>
</configuration>
EOF
$ cat > $HADOOP_CONF_DIR/hdfs-site.xml <<EOF
<configuration>
<!-- 1, not 3: there is only one node -->
<property><name>dfs.replication</name><value>1</value></property>
<property><name>dfs.namenode.name.dir</name><value>file://$HOME/hadoop_store/hdfs/namenode</value></property>
<property><name>dfs.datanode.data.dir</name><value>file://$HOME/hadoop_store/hdfs/datanode</value></property>
</configuration>
EOF
$ cat > $HADOOP_CONF_DIR/mapred-site.xml <<EOF
<configuration>
<property><name>mapreduce.framework.name</name><value>yarn</value></property>
<property><name>yarn.app.mapreduce.am.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
<property><name>mapreduce.map.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
<property><name>mapreduce.reduce.env</name><value>HADOOP_MAPRED_HOME=$HADOOP_HOME</value></property>
</configuration>
EOF
$ cat > $HADOOP_CONF_DIR/yarn-site.xml <<'EOF'
<configuration>
<property><name>yarn.nodemanager.aux-services</name><value>mapreduce_shuffle</value></property>
</configuration>
EOF
$ echo "export JAVA_HOME=$JAVA_HOME" >> $HADOOP_CONF_DIR/hadoop-env.sh
$ hdfs namenode -format 2>&1 | grep "has been successfully formatted"
2026-10-04 23:03:54,977 INFO common.Storage: Storage directory ~/hadoop_store/hdfs/namenode has been successfully formatted.
$ hdfs --daemon start namenode
$ hdfs --daemon start datanode
$ hdfs --daemon start secondarynamenode
$ yarn --daemon start resourcemanager
$ yarn --daemon start nodemanager
$ sleep 15
$ jps | awk '{print $2}' | sort
DataNode
Jps
NameNode
NodeManager
ResourceManager
SecondaryNameNode
$ hdfs dfsadmin -report 2>/dev/null | grep "Live datanodes"
Live datanodes (1):
$ hdfs dfs -mkdir -p /user/$USER && hdfs dfs -ls /user
Found 1 items
drwxr-xr-x - root supergroup 0 2026-10-04 23:04 /user/root
$ curl -sL -o /dev/null -w "NameNode http://localhost:9870 %{http_code}\n" http://localhost:9870
NameNode http://localhost:9870 200
$ curl -s -o /dev/null -w "ResourceManager http://localhost:8088 %{http_code}\n" http://localhost:8088/cluster
ResourceManager http://localhost:8088 200
$ yarn --daemon stop nodemanager
$ yarn --daemon stop resourcemanager
$ hdfs --daemon stop secondarynamenode
$ hdfs --daemon stop datanode
$ hdfs --daemon stop namenode
This one runs alone, not on the lab's cluster: _drive_01_install_hadoop.py gives it an
empty home directory and the Hadoop tarball, and the script does everything else — unpacks,
configures, formats, starts, checks and stops. Where it was run there is no ssh server, so the
daemons are started one by one with hdfs --daemon start, the command start-dfs.sh runs on
each host over ssh.
RESULT
Hadoop 3.3.6 installed in an empty home directory, HDFS formatted, the five daemons running — jps lists all five — one live DataNode, and both web interfaces answering 200. Two corrections: the download moved to archive.apache.org, and MapReduce needs HADOOP_MAPRED_HOME passed to its containers.
Explore the Hadoop directory structure and its basic commands — the hadoop fs operations.
Make, list, copy, move and remove files in HDFS; count the space they use; set permissions and replication; and ask where a file's blocks are.
On the cluster, 02_hdfs_commands.sh:
THE FOUR HADOOP FS FACTS WORTH MARKS
hadoop fs and hdfs dfs are the same command. hadoop fs also works
on local and S3 paths.
There is no cd. HDFS has no working directory — every path is
absolute or relative to /user/$USER.
-rm moves to .Trash and still costs quota for a day.
On the cluster, 02_hdfs_commands.sh:
# Experiment 2 -- explore the Hadoop directory structure and basic hadoop fs commands
#
# Run it: bash 02_hdfs_commands.sh, with a cluster running (experiment 1). It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is none -- these are filesystem commands, verified by the arithmetic in 04
#
# Step 1: Make a directory, and list it
hdfs dfs -mkdir -p /user/student/sales
hdfs dfs -ls /user/student
hdfs dfs -ls -R /user # recursive
# Step 2: Put a file in, and read it back
hdfs dfs -put sales.csv /user/student/sales/
hdfs dfs -cat /user/student/sales/sales.csv | head
hdfs dfs -tail /user/student/sales/sales.csv
hdfs dfs -get /user/student/sales/sales.csv ./back.csv
# Step 3: Copy, move and remove
hdfs dfs -cp /user/student/sales/sales.csv /user/student/copy.csv
hdfs dfs -mv /user/student/copy.csv /user/student/moved.csv
hdfs dfs -rm -r /user/student/sales # goes to .Trash, not to nothing
# Step 4: Count the space used
hdfs dfs -du -h /user/student
hdfs dfs -df -h /
hdfs dfs -count /user/student # DIRS FILES BYTES
# Step 5: Set permissions and replication
hdfs dfs -chmod 640 /user/student/moved.csv
hdfs dfs -chown student:analysts /user/student/moved.csv
hdfs dfs -setrep -w 2 /user/student/moved.csv # change replication
hdfs dfs -stat "%r %o %b" /user/student/moved.csv # replication blocksize bytes
# Step 6: Ask where the blocks are, and how the cluster is
hdfs fsck /user/student -files -blocks -locations # WHERE each block lives
hdfs dfsadmin -report # per-DataNode capacity
hdfs dfsadmin -safemode get # ON during startup
# --- the four things that surprise people ----------------------------------
# 1. `hadoop fs` and `hdfs dfs` are the same command. `hadoop fs` also works
# on local and S3 paths; `hdfs dfs` is HDFS only.
# 2. THERE IS NO `cd`. HDFS has no working directory -- every path is
# absolute, or relative to /user/$USER.
# 3. `-rm` moves to .Trash and still costs quota for a day. Use -skipTrash
# when you mean it.
# 4. There is no in-place edit. HDFS is WRITE-ONCE, APPEND-ONLY: to change one
# byte you rewrite the file. That single constraint is why HDFS can drop
# file locking, and why it suits analytics and not OLTP.
On the cluster, 02_hdfs_commands.sh:
OUTPUT
$ hdfs dfs -mkdir -p /user/student/sales
$ hdfs dfs -ls /user/student
Found 1 items
drwxr-xr-x - root supergroup 0 2026-10-04 23:05 /user/student/sales
$ hdfs dfs -ls -R /user
drwxr-xr-x - root supergroup 0 2026-10-04 23:05 /user/student
drwxr-xr-x - root supergroup 0 2026-10-04 23:05 /user/student/sales
$ hdfs dfs -put sales.csv /user/student/sales/
$ hdfs dfs -cat /user/student/sales/sales.csv | head
date_key,store,region,product,category,qty,list_price
D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0
D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0
D1,Guntur,South,Tea 500g,Grocery,8,210.0
D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0
D2,Hyderabad,North,Notebook,Stationery,20,40.0
D3,Guntur,South,Tea 500g,Grocery,12,210.0
D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0
D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0
D4,Hyderabad,North,Notebook,Stationery,15,40.0
$ hdfs dfs -tail /user/student/sales/sales.csv
date_key,store,region,product,category,qty,list_price
D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0
D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0
D1,Guntur,South,Tea 500g,Grocery,8,210.0
D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0
D2,Hyderabad,North,Notebook,Stationery,20,40.0
D3,Guntur,South,Tea 500g,Grocery,12,210.0
D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0
D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0
D4,Hyderabad,North,Notebook,Stationery,15,40.0
$ hdfs dfs -get /user/student/sales/sales.csv ./back.csv
$ hdfs dfs -cp /user/student/sales/sales.csv /user/student/copy.csv
$ hdfs dfs -mv /user/student/copy.csv /user/student/moved.csv
$ hdfs dfs -rm -r /user/student/sales
2026-10-04 23:06:06,486 INFO fs.TrashPolicyDefault: Moved: 'hdfs://localhost:9000/user/student/sales' to trash at: hdfs://localhost:9000/user/root/.Trash/Current/user/student/sales
$ hdfs dfs -du -h /user/student
468 1.4 K /user/student/moved.csv
$ hdfs dfs -df -h /
Filesystem Size Used Available Use%
hdfs://localhost:9000 1007.9 G 98.8 K 36.1 G 0%
$ hdfs dfs -count /user/student
1 1 468 /user/student
$ hdfs dfs -chmod 640 /user/student/moved.csv
$ hdfs dfs -chown student:analysts /user/student/moved.csv
$ hdfs dfs -setrep -w 2 /user/student/moved.csv
Replication 2 set: /user/student/moved.csv
Waiting for /user/student/moved.csv ...
WARNING: the waiting time may be long for DECREASING the number of replications.
. done
$ hdfs dfs -stat "%r %o %b" /user/student/moved.csv
2 134217728 468
$ hdfs fsck /user/student -files -blocks -locations
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&locations=1&path=%2Fuser%2Fstudent
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student at Sun Oct 04 23:06:30 UTC 2026
/user/student <dir>
/user/student/moved.csv 468 bytes, replicated: replication=2, 1 block(s): OK
0. BP-1985187428-127.0.0.1-1791155116848:blk_1073741826_1002 len=468 Live_repl=2 [DatanodeInfoWithStorage[127.0.0.1:9886,DS-aeb1c4f1-1445-4508-8b83-5f8852f1b202,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-c7bba049-11ba-4c05-a5e4-c385376f29cb,DISK]]
Status: HEALTHY
Number of data-nodes: 4
Number of racks: 1
Total dirs: 1
Total symlinks: 0
Replicated Blocks:
Total size: 468 B
Total files: 1
Total blocks (validated): 1 (avg. block size 468 B)
Minimally replicated blocks: 1 (100.0 %)
Over-replicated blocks: 0 (0.0 %)
Under-replicated blocks: 0 (0.0 %)
Mis-replicated blocks: 0 (0.0 %)
Default replication factor: 3
Average block replication: 2.0
Missing blocks: 0
Corrupt blocks: 0
Missing replicas: 0 (0.0 %)
Blocks queued for replication: 0
Erasure Coded Block Groups:
Total size: 0 B
Total files: 0
Total block groups (validated): 0
Minimally erasure-coded block groups: 0
Over-erasure-coded block groups: 0
Under-erasure-coded block groups: 0
Unsatisfactory placement block groups: 0
Average block group size: 0.0
Missing block groups: 0
Corrupt block groups: 0
Missing internal blocks: 0
Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:06:30 UTC 2026 in 9 milliseconds
The filesystem under path '/user/student' is HEALTHY
$ hdfs dfsadmin -report
Configured Capacity: 1082212696064 (1007.89 GB)
Present Capacity: 38770710875 (36.11 GB)
DFS Remaining: 38770610176 (36.11 GB)
DFS Used: 100699 (98.34 KB)
DFS Used%: 0.00%
Replicated Blocks:
Under replicated blocks: 0
Blocks with corrupt replicas: 0
Missing blocks: 0
Missing blocks (with replication factor 1): 0
Low redundancy blocks with highest priority to recover: 0
Pending deletion blocks: 0
Erasure Coded Block Groups:
Low redundancy block groups: 0
Block groups with corrupt internal blocks: 0
Missing block groups: 0
Low redundancy blocks with highest priority to recover: 0
Pending deletion blocks: 0
-------------------------------------------------
Live datanodes (4):
Name: 127.0.0.1:9866 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088212001 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:22 UTC 2026
Num of Blocks: 1
Name: 127.0.0.1:9886 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088212001 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:25 UTC 2026
Num of Blocks: 1
Name: 127.0.0.1:9896 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25534 (24.94 KB)
Non DFS Used: 30088211522 (28.02 GB)
DFS Remaining: 9692651520 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:31 UTC 2026
Last Block Report: Sun Oct 04 23:05:28 UTC 2026
Num of Blocks: 2
Name: 127.0.0.1:9906 (localhost)
Hostname: localhost
Decommission Status : Normal
Configured Capacity: 270553174016 (251.97 GB)
DFS Used: 25055 (24.47 KB)
Non DFS Used: 30088207905 (28.02 GB)
DFS Remaining: 9692655616 (9.03 GB)
DFS Used%: 0.00%
DFS Remaining%: 3.58%
Configured Cache Capacity: 0 (0 B)
Cache Used: 0 (0 B)
Cache Remaining: 0 (0 B)
Cache Used%: 100.00%
Cache Remaining%: 0.00%
Xceivers: 0
Last contact: Sun Oct 04 23:06:30 UTC 2026
Last Block Report: Sun Oct 04 23:05:30 UTC 2026
Num of Blocks: 1
$ hdfs dfsadmin -safemode get
Safe mode is OFF
RESULT
Every hdfs dfs operation ran against the cluster. A removed file goes to .Trash, and fsck and dfsadmin -report show where each block lives and what each DataNode holds.
Demonstrate the Hadoop architecture components — HDFS, YARN and MapReduce — using sample logs.
Run one MapReduce job over a sample log, then follow it through the logs of the daemons that ran it.
On the cluster, 03_architecture.sh:
THE LOG TRACE TO FOLLOW
RM: application submitted → NM: AM container started → RM: map containers assigned → NM: map tasks start → NameNode: block reads served locally → RM: reduce containers assigned → NM: reduce fetches map outputs, over HTTP → RM: SUCCEEDED.
The one line worth finding is the shuffle fetch. It is the only step where data crosses the network in bulk, and it is what the combiner in experiment 7 exists to shrink.
On the cluster, 03_architecture.sh:
# Experiment 3 -- demonstrate the Hadoop architecture components using sample logs
#
# Run it: bash 03_architecture.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 06_yarn_scheduling.py, which runs the scheduler the logs describe
#
# Step 1: List the daemons
jps | awk '{print $2}' | sort # the daemons, one JVM each
# Step 2: Run a job, and find it
hdfs dfs -mkdir -p /user/student/logs
hdfs dfs -put syslog /user/student/logs/
yarn jar $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar \
wordcount /user/student/logs /user/student/wc-out 2>&1 | grep -E "Running job|completed successfully"
# [Corrected: this put /var/log/syslog, which many systems no longer have
# (journald keeps the log instead), into a directory that did not exist yet.
# Any text file will do; syslog here is a few lines of the lab's documents.]
APP=$(yarn application -list -appStates ALL 2>/dev/null | awk '/^application_/{print $1}' | tail -1)
yarn application -list -appStates ALL 2>/dev/null | awk -F'\t' 'NR>1{print $1 " | " $2 " | " $6 " | " $7}'
yarn application -status $APP 2>/dev/null | grep -E "Application-Name|State|Final-State"
sleep 10 # the NodeManager uploads the logs when the job ends
yarn logs -applicationId $APP 2>/dev/null | grep "^Container: " | sort -u | sed -E 's/ on .*//'
# [Corrected: the commands used application_1699999999999_0001, an id from
# someone else's cluster. Every cluster numbers its own; take it from
# `yarn application -list`.]
# Step 3: Read each daemon's log
# [Corrected: these were `tail -f`, which never returns -- each would have to
# be stopped with Ctrl-C before the next. grep finds the lines described. And
# Hadoop 3 names the YARN logs hadoop-<user>-resourcemanager-<host>.log, not yarn-.]
grep -h -m 2 "BLOCK\* allocate" $HADOOP_LOG_DIR/hadoop-*-namenode-*.log | cut -c 25-
# BlockStateChange lines: every allocation and replication decision
grep -h -m 1 "Receiving BP-" $HADOOP_LOG_DIR/hadoop-*-datanode-*.log | cut -c 25-
# "Receiving BP-...:blk_..." -- a block landing, with its pipeline
grep -h -m 2 "Assigned container container_" $HADOOP_LOG_DIR/*-resourcemanager-*.log | cut -c 25-
# "Assigned container container_..." -- the scheduler, deciding
grep -h -m 1 "Starting resource-monitoring for container_" $HADOOP_LOG_DIR/*-nodemanager-*.log | cut -c 25-
# "Starting resource-monitoring for container_..." -- the container's life
yarn logs -applicationId $APP 2>/dev/null | grep -m 1 -oE "fetcher#[0-9]+ about to shuffle output of map [^ ]+"
# in the reduce container's own log: one of the reduce's fetchers asks the
# NodeManager's shuffle service, over HTTP, for a map's output -- THE SHUFFLE
# --- the trace to follow, in order -----------------------------------------
# RM log : application submitted, ApplicationMaster container assigned
# NM log : AM container started
# RM log : AM requests N map containers; scheduler assigns them
# NM logs : each map task starts, reports progress
# NameNode log : block reads served, LOCAL where possible
# RM log : reduce containers assigned after map progress passes 5%
# NM log : reduce fetches map outputs -- THE SHUFFLE, over HTTP
# RM log : application FINISHED, SUCCEEDED
#
# The one line worth finding is the shuffle fetch. It is the only step where
# data crosses the network in bulk, and it is what the combiner in
# experiment 7 exists to shrink.
On the cluster, 03_architecture.sh:
OUTPUT
$ jps | awk '{print $2}' | sort
DataNode
DataNode
DataNode
DataNode
JobHistoryServer
Jps
NameNode
NodeManager
ResourceManager
SecondaryNameNode
$ hdfs dfs -mkdir -p /user/student/logs
$ hdfs dfs -put syslog /user/student/logs/
$ yarn jar $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar \
wordcount /user/student/logs /user/student/wc-out 2>&1 | grep -E "Running job|completed successfully"
2026-10-04 23:40:14,246 INFO mapreduce.Job: Running job: job_1791157202192_0001
2026-10-04 23:40:31,543 INFO mapreduce.Job: Job job_1791157202192_0001 completed successfully
$ APP=$(yarn application -list -appStates ALL 2>/dev/null | awk '/^application_/{print $1}' | tail -1)
$ yarn application -list -appStates ALL 2>/dev/null | awk -F'\t' 'NR>1{print $1 " | " $2 " | " $6 " | " $7}'
Application-Id | Application-Name | State | Final-State
application_1791157202192_0001 | word count | FINISHED | SUCCEEDED
$ yarn application -status $APP 2>/dev/null | grep -E "Application-Name|State|Final-State"
Application-Name : word count
State : FINISHED
Final-State : SUCCEEDED
$ sleep 10
$ yarn logs -applicationId $APP 2>/dev/null | grep "^Container: " | sort -u | sed -E 's/ on .*//'
Container: container_1791157202192_0001_01_000001
Container: container_1791157202192_0001_01_000002
Container: container_1791157202192_0001_01_000003
$ grep -h -m 2 "BLOCK\* allocate" $HADOOP_LOG_DIR/hadoop-*-namenode-*.log | cut -c 25-
INFO org.apache.hadoop.hdfs.StateChange: BLOCK* allocate blk_1073741825_1001, replicas=127.0.0.1:9886, 127.0.0.1:9866, 127.0.0.1:9896 for /user/student/logs/syslog._COPYING_
INFO org.apache.hadoop.hdfs.StateChange: BLOCK* allocate blk_1073741826_1002, replicas=127.0.0.1:9886, 127.0.0.1:9896, 127.0.0.1:9906 for /tmp/hadoop-yarn/staging/root/.staging/job_1791157202192_0001/job.jar
$ grep -h -m 1 "Receiving BP-" $HADOOP_LOG_DIR/hadoop-*-datanode-*.log | cut -c 25-
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:57158 dest: /127.0.0.1:9886
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:35392 dest: /127.0.0.1:9896
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741826_1002 src: /127.0.0.1:32904 dest: /127.0.0.1:9906
INFO org.apache.hadoop.hdfs.server.datanode.DataNode: Receiving BP-1713403942-127.0.0.1-1791157181747:blk_1073741825_1001 src: /127.0.0.1:42626 dest: /127.0.0.1:9866
$ grep -h -m 2 "Assigned container container_" $HADOOP_LOG_DIR/*-resourcemanager-*.log | cut -c 25-
INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.common.fica.FiCaSchedulerNode: Assigned container container_1791157202192_0001_01_000001 of capacity <memory:512, vCores:1> on host localhost:40995, which has 1 containers, <memory:512, vCores:1> used and <memory:5632, vCores:3> available after allocation
INFO org.apache.hadoop.yarn.server.resourcemanager.scheduler.common.fica.FiCaSchedulerNode: Assigned container container_1791157202192_0001_01_000002 of capacity <memory:512, vCores:1> on host localhost:40995, which has 2 containers, <memory:1024, vCores:2> used and <memory:5120, vCores:2> available after allocation
$ grep -h -m 1 "Starting resource-monitoring for container_" $HADOOP_LOG_DIR/*-nodemanager-*.log | cut -c 25-
INFO org.apache.hadoop.yarn.server.nodemanager.containermanager.monitor.ContainersMonitorImpl: Starting resource-monitoring for container_1791157202192_0001_01_000001
$ yarn logs -applicationId $APP 2>/dev/null | grep -m 1 -oE "fetcher#[0-9]+ about to shuffle output of map [^ ]+"
fetcher#2 about to shuffle output of map attempt_1791157202192_0001_m_000000_0
RESULT
One job, traced through four daemons: the ResourceManager admitted it and allocated its containers, the NodeManager launched them, the NameNode served the blocks, and the reduce fetched the map outputs over HTTP — the shuffle.
Store and retrieve a large file in HDFS, and demonstrate its block distribution and replication factor.
Put a 300 MB file into HDFS, find its blocks and where each replica is, change its replication, and work out the block arithmetic and the small-files cost.
On the cluster, 04_hdfs_store.sh:
The Python check, 04_blocks_replication.py:
THE BLOCK TABLE
| File | Blocks | Last block | Disk used |
|---|---|---|---|
| 1 MB | 1 | 1.00 MB | 1 MB |
| 128 MB | 1 | 128.00 MB | 128 MB |
| 129 MB | 2 | 1.00 MB | 129 MB |
| 260 MB | 3 | 4.00 MB | 260 MB |
| 1024 MB | 8 | 128.00 MB | 1024 MB |
| 5000 MB | 40 | 8.00 MB | 5000 MB |
A 260 MB file is 128 + 128 + 4. HDFS wastes no space on block padding, and this is the most examined calculation in the course.
But 128 MB → 1 block and 129 MB → 2. One byte past the boundary costs a whole block object in NameNode RAM, though almost no disk.
On the cluster, 04_hdfs_store.sh:
# Experiment 4 -- store and retrieve large files in HDFS -- block distribution and replication
#
# Run it: bash 04_hdfs_store.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 04_blocks_replication.py, which computes every figure below
#
# make a file bigger than one block so there is something to distribute
# Step 1: Make a 300 MB file
dd if=/dev/urandom of=big.bin bs=1M count=300 # 300 MB -> 3 blocks
# Step 2: Put it in HDFS
hdfs dfs -mkdir -p /user/student/big
hdfs dfs -put big.bin /user/student/big/
# how many blocks, and where are they?
# Step 3: Find its blocks and their replicas
hdfs fsck /user/student/big/big.bin -files -blocks -locations
# expect: 3 blocks -- 128 MB, 128 MB, 44 MB
# the LAST BLOCK IS SHORT. HDFS does not pad.
# Step 4: Change its replication
hdfs dfs -stat "%r" /user/student/big/big.bin # replication factor
hdfs dfs -setrep -w 2 /user/student/big/big.bin # -w waits for completion
hdfs fsck /user/student/big/big.bin -files -blocks # now 2 locations per block
# a non-default block size, set PER FILE at write time
# Step 5: Write it again with 64 MB blocks
hdfs dfs -D dfs.blocksize=67108864 -put big.bin /user/student/big/small-blocks.bin
hdfs fsck /user/student/big/small-blocks.bin -files -blocks
# expect: 5 blocks -- four of 64 MB and a last one of 44 MB -- more blocks,
# more NameNode objects, more map tasks (one per block by default)
# [Corrected: this said "5 blocks of 64 MB". 300 MB is 4 x 64 + 44: the last
# block is short here too, as it is above.]
# retrieve and verify
# Step 6: Get it back, and compare
hdfs dfs -get /user/student/big/big.bin ./back.bin
md5sum big.bin back.bin # must match
hdfs dfsadmin -report | grep -E "Name|DFS Used|Remaining"
The Python check, 04_blocks_replication.py:
"""Experiment 4 -- store and retrieve a large file in HDFS: blocks, block
distribution and the replication factor.
`04_hdfs_store.sh` carries the commands you actually type, and runs on a real
cluster (the lab page shows it). What runs here is the ARITHMETIC, which is the
part that gets examined and the part students get wrong.
"""
from blocks import BLOCK, MB, blocks_for, namenode_memory, placement
def main():
print(" Experiment 4 -- HDFS blocks, distribution and replication")
# Step 1: Size the blocks
print("\n block sizing (default block = 128 MB):")
print(f" {'file':>10} {'blocks':>6} {'last block':>12} {'disk used':>12}")
for mb in (1, 128, 129, 260, 1024, 5000):
n, last = blocks_for(mb * MB)
print(f" {mb:>7} MB {n:>6} {last / MB:>9.2f} MB {mb:>9} MB")
print(""" a 260 MB file is 128 + 128 + 4, NOT three full blocks.
An HDFS block is a logical MAXIMUM; the last block occupies
only what it needs. HDFS wastes no space on block padding --
which is the opposite of what the word 'block' suggests, and
the mistake to avoid in the exam""")
n1, last1 = blocks_for(1 * MB)
assert (n1, last1) == (1, 1 * MB)
n260, last260 = blocks_for(260 * MB)
assert n260 == 3 and last260 == 4 * MB
n128, _ = blocks_for(128 * MB)
n129, _ = blocks_for(129 * MB)
assert (n128, n129) == (1, 2), "one byte over a block boundary costs a block"
print("\n 128 MB -> 1 block, 129 MB -> 2 blocks")
print(""" one byte past the boundary costs a whole block OBJECT
in NameNode memory, though almost no disk. Metadata is the
scarce resource in HDFS, not disk""")
# Step 2: Place a 1 GB file's replicas
print("\n a 1 GB file, replication 3, 6 DataNodes across 2 racks:")
n, last = blocks_for(1024 * MB)
plan, rack_of = placement(n, 3, datanodes=6, racks=2)
print(f" {n} blocks (last = {last / MB:.0f} MB)")
print(f" {'block':>6} {'replicas (node/rack)':<34} racks used")
for i, nodes in enumerate(plan):
desc = " ".join(f"n{d}/r{rack_of[d]}" for d in nodes)
print(f" {i:>6} {desc:<34} {len({rack_of[d] for d in nodes})}")
for nodes in plan:
assert len({rack_of[d] for d in nodes}) == 2, "must span two racks"
assert len(set(nodes)) == 3, "three replicas on three distinct nodes"
print(""" every block spans EXACTLY TWO racks: one replica on the
writer's rack, two on another. Two racks survive a rack
failure; a third rack would double cross-rack write traffic
to buy very little. That trade is the whole policy""")
raw = 1024
print(f"\n storage cost: {raw} MB of data at replication 3 "
f"occupies {raw * 3} MB of disk")
print(f" the same data with erasure coding (RS-6-3) would occupy "
f"{raw * 9 // 6} MB")
assert raw * 3 == 3072 and raw * 9 // 6 == 1536
print(""" replication costs 200% overhead for 3x durability;
RS-6-3 erasure coding costs 50% for comparable durability,
at the price of expensive reconstruction reads. HDFS added
erasure coding in 3.0 for exactly this reason -- COLD data""")
# Step 3: Count the small-files cost
print("\n the small-files problem, in NameNode RAM (~150 bytes/object):")
print(f" {'scenario':<26}{'files':>12}{'blocks':>12}{'NameNode RAM':>16}")
for label, files, size_each in (
("one 1 GB file", 1, 1024 * MB),
("1,000 x 1 MB files", 1000, 1 * MB),
("1,000,000 x 1 KB files", 1_000_000, 1024)):
total_blocks = sum(blocks_for(size_each)[0] for _ in range(1)) * files
ram = namenode_memory(files, total_blocks)
print(f" {label:<26}{files:>12,}{total_blocks:>12,}"
f"{ram / MB:>13.2f} MB")
one = namenode_memory(1, 8)
many = namenode_memory(1_000_000, 1_000_000)
assert many // one > 200_000
print(f""" the same 1 GB costs {one} bytes as one file and
{many / MB:.0f} MB as a million small ones -- a factor of
{many // one:,}. HDFS was built for few large files, and this
single table is the reason""")
if __name__ == "__main__":
main()
On the cluster, 04_hdfs_store.sh:
OUTPUT
$ dd if=/dev/urandom of=big.bin bs=1M count=300
300+0 records in
300+0 records out
314572800 bytes (315 MB, 300 MiB) copied, 2.09212 s, 150 MB/s
$ hdfs dfs -mkdir -p /user/student/big
$ hdfs dfs -put big.bin /user/student/big/
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&locations=1&path=%2Fuser%2Fstudent%2Fbig%2Fbig.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/big.bin at Sun Oct 04 23:08:33 UTC 2026
/user/student/big/big.bin 314572800 bytes, replicated: replication=3, 3 block(s): OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741825_1001 len=134217728 Live_repl=3 [DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-1431a54e-afc3-4c6b-bc5f-702712a1ce03,DISK]]
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741826_1002 len=134217728 Live_repl=3 [DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9896,DS-1431a54e-afc3-4c6b-bc5f-702712a1ce03,DISK]]
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741827_1003 len=46137344 Live_repl=3 [DatanodeInfoWithStorage[127.0.0.1:9906,DS-dd89703a-f511-44f9-9d22-8afb11511b2c,DISK], DatanodeInfoWithStorage[127.0.0.1:9866,DS-3739b86c-18ff-4fdc-8ba5-0d9039f4a9aa,DISK], DatanodeInfoWithStorage[127.0.0.1:9886,DS-2d1746a3-8235-40de-befd-514261a60ca2,DISK]]
Status: HEALTHY
Number of data-nodes: 4
Number of racks: 1
Total dirs: 0
Total symlinks: 0
Replicated Blocks:
Total size: 314572800 B
Total files: 1
Total blocks (validated): 3 (avg. block size 104857600 B)
Minimally replicated blocks: 3 (100.0 %)
Over-replicated blocks: 0 (0.0 %)
Under-replicated blocks: 0 (0.0 %)
Mis-replicated blocks: 0 (0.0 %)
Default replication factor: 3
Average block replication: 3.0
Missing blocks: 0
Corrupt blocks: 0
Missing replicas: 0 (0.0 %)
Blocks queued for replication: 0
Erasure Coded Block Groups:
Total size: 0 B
Total files: 0
Total block groups (validated): 0
Minimally erasure-coded block groups: 0
Over-erasure-coded block groups: 0
Under-erasure-coded block groups: 0
Unsatisfactory placement block groups: 0
Average block group size: 0.0
Missing block groups: 0
Corrupt block groups: 0
Missing internal blocks: 0
Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:33 UTC 2026 in 10 milliseconds
The filesystem under path '/user/student/big/big.bin' is HEALTHY
$ hdfs dfs -stat "%r" /user/student/big/big.bin
3
$ hdfs dfs -setrep -w 2 /user/student/big/big.bin
Replication 2 set: /user/student/big/big.bin
Waiting for /user/student/big/big.bin ...
WARNING: the waiting time may be long for DECREASING the number of replications.
. done
$ hdfs fsck /user/student/big/big.bin -files -blocks
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&path=%2Fuser%2Fstudent%2Fbig%2Fbig.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/big.bin at Sun Oct 04 23:08:48 UTC 2026
/user/student/big/big.bin 314572800 bytes, replicated: replication=2, 3 block(s): OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741825_1001 len=134217728 Live_repl=2
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741826_1002 len=134217728 Live_repl=2
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741827_1003 len=46137344 Live_repl=2
Status: HEALTHY
Number of data-nodes: 4
Number of racks: 1
Total dirs: 0
Total symlinks: 0
Replicated Blocks:
Total size: 314572800 B
Total files: 1
Total blocks (validated): 3 (avg. block size 104857600 B)
Minimally replicated blocks: 3 (100.0 %)
Over-replicated blocks: 0 (0.0 %)
Under-replicated blocks: 0 (0.0 %)
Mis-replicated blocks: 0 (0.0 %)
Default replication factor: 3
Average block replication: 2.0
Missing blocks: 0
Corrupt blocks: 0
Missing replicas: 0 (0.0 %)
Blocks queued for replication: 0
Erasure Coded Block Groups:
Total size: 0 B
Total files: 0
Total block groups (validated): 0
Minimally erasure-coded block groups: 0
Over-erasure-coded block groups: 0
Under-erasure-coded block groups: 0
Unsatisfactory placement block groups: 0
Average block group size: 0.0
Missing block groups: 0
Corrupt block groups: 0
Missing internal blocks: 0
Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:48 UTC 2026 in 1 milliseconds
The filesystem under path '/user/student/big/big.bin' is HEALTHY
$ hdfs dfs -D dfs.blocksize=67108864 -put big.bin /user/student/big/small-blocks.bin
$ hdfs fsck /user/student/big/small-blocks.bin -files -blocks
Connecting to namenode via http://localhost:9870/fsck?ugi=root&files=1&blocks=1&path=%2Fuser%2Fstudent%2Fbig%2Fsmall-blocks.bin
FSCK started by root (auth:SIMPLE) from /127.0.0.1 for path /user/student/big/small-blocks.bin at Sun Oct 04 23:08:55 UTC 2026
/user/student/big/small-blocks.bin 314572800 bytes, replicated: replication=3, 5 block(s): OK
0. BP-245208533-127.0.0.1-1791155276951:blk_1073741828_1004 len=67108864 Live_repl=3
1. BP-245208533-127.0.0.1-1791155276951:blk_1073741829_1005 len=67108864 Live_repl=3
2. BP-245208533-127.0.0.1-1791155276951:blk_1073741830_1006 len=67108864 Live_repl=3
3. BP-245208533-127.0.0.1-1791155276951:blk_1073741831_1007 len=67108864 Live_repl=3
4. BP-245208533-127.0.0.1-1791155276951:blk_1073741832_1008 len=46137344 Live_repl=3
Status: HEALTHY
Number of data-nodes: 4
Number of racks: 1
Total dirs: 0
Total symlinks: 0
Replicated Blocks:
Total size: 314572800 B
Total files: 1
Total blocks (validated): 5 (avg. block size 62914560 B)
Minimally replicated blocks: 5 (100.0 %)
Over-replicated blocks: 0 (0.0 %)
Under-replicated blocks: 0 (0.0 %)
Mis-replicated blocks: 0 (0.0 %)
Default replication factor: 3
Average block replication: 3.0
Missing blocks: 0
Corrupt blocks: 0
Missing replicas: 0 (0.0 %)
Blocks queued for replication: 0
Erasure Coded Block Groups:
Total size: 0 B
Total files: 0
Total block groups (validated): 0
Minimally erasure-coded block groups: 0
Over-erasure-coded block groups: 0
Under-erasure-coded block groups: 0
Unsatisfactory placement block groups: 0
Average block group size: 0.0
Missing block groups: 0
Corrupt block groups: 0
Missing internal blocks: 0
Blocks queued for replication: 0
FSCK ended at Sun Oct 04 23:08:55 UTC 2026 in 2 milliseconds
The filesystem under path '/user/student/big/small-blocks.bin' is HEALTHY
$ hdfs dfs -get /user/student/big/big.bin ./back.bin
$ md5sum big.bin back.bin
6792d150c0ab35807e5020be57771bb5 big.bin
6792d150c0ab35807e5020be57771bb5 back.bin
$ hdfs dfsadmin -report | grep -E "Name|DFS Used|Remaining"
DFS Remaining: 29910220800 (27.86 GB)
DFS Used: 1585250451 (1.48 GB)
DFS Used%: 5.03%
Name: 127.0.0.1:9866 (localhost)
DFS Used: 498819114 (475.71 MB)
Non DFS Used: 31804215254 (29.62 GB)
DFS Remaining: 7477854208 (6.96 GB)
DFS Used%: 0.18%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9886 (localhost)
DFS Used: 181788693 (173.37 MB)
Non DFS Used: 32121655275 (29.92 GB)
DFS Remaining: 7477444608 (6.96 GB)
DFS Used%: 0.07%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9896 (localhost)
DFS Used: 587587633 (560.37 MB)
Non DFS Used: 31715823567 (29.54 GB)
DFS Remaining: 7477477376 (6.96 GB)
DFS Used%: 0.22%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
Name: 127.0.0.1:9906 (localhost)
DFS Used: 317055011 (302.37 MB)
Non DFS Used: 31986388957 (29.79 GB)
DFS Remaining: 7477444608 (6.96 GB)
DFS Used%: 0.12%
DFS Remaining%: 2.76%
Cache Remaining: 0 (0 B)
Cache Remaining%: 0.00%
The Python check, 04_blocks_replication.py:
OUTPUT
Experiment 4 -- HDFS blocks, distribution and replication
block sizing (default block = 128 MB):
file blocks last block disk used
1 MB 1 1.00 MB 1 MB
128 MB 1 128.00 MB 128 MB
129 MB 2 1.00 MB 129 MB
260 MB 3 4.00 MB 260 MB
1024 MB 8 128.00 MB 1024 MB
5000 MB 40 8.00 MB 5000 MB
a 260 MB file is 128 + 128 + 4, NOT three full blocks.
An HDFS block is a logical MAXIMUM; the last block occupies
only what it needs. HDFS wastes no space on block padding --
which is the opposite of what the word 'block' suggests, and
the mistake to avoid in the exam
128 MB -> 1 block, 129 MB -> 2 blocks
one byte past the boundary costs a whole block OBJECT
in NameNode memory, though almost no disk. Metadata is the
scarce resource in HDFS, not disk
a 1 GB file, replication 3, 6 DataNodes across 2 racks:
8 blocks (last = 128 MB)
block replicas (node/rack) racks used
0 n0/r0 n1/r1 n3/r1 2
1 n2/r0 n3/r1 n5/r1 2
2 n4/r0 n5/r1 n1/r1 2
3 n0/r0 n1/r1 n5/r1 2
4 n2/r0 n3/r1 n1/r1 2
5 n4/r0 n5/r1 n3/r1 2
6 n0/r0 n1/r1 n3/r1 2
7 n2/r0 n3/r1 n5/r1 2
every block spans EXACTLY TWO racks: one replica on the
writer's rack, two on another. Two racks survive a rack
failure; a third rack would double cross-rack write traffic
to buy very little. That trade is the whole policy
storage cost: 1024 MB of data at replication 3 occupies 3072 MB of disk
the same data with erasure coding (RS-6-3) would occupy 1536 MB
replication costs 200% overhead for 3x durability;
RS-6-3 erasure coding costs 50% for comparable durability,
at the price of expensive reconstruction reads. HDFS added
erasure coding in 3.0 for exactly this reason -- COLD data
the small-files problem, in NameNode RAM (~150 bytes/object):
scenario files blocks NameNode RAM
one 1 GB file 1 8 0.00 MB
1,000 x 1 MB files 1,000 1,000 0.29 MB
1,000,000 x 1 KB files 1,000,000 1,000,000 286.10 MB
the same 1 GB costs 1350 bytes as one file and
286 MB as a million small ones -- a factor of
222,222. HDFS was built for few large files, and this
single table is the reason
REPLICA PLACEMENT, 1 GB OVER 6 NODES IN 2 RACKS
| Block | Replicas | Racks |
|---|---|---|
| 0 | n0/r0, n1/r1, n3/r1 | 2 |
| 1 | n2/r0, n3/r1, n5/r1 | 2 |
| 2 | n4/r0, n5/r1, n1/r1 | 2 |
| … | … | 2 |
Every block spans exactly two racks — one replica on the writer's rack, two on another. Asserted for all 8 blocks.
THE SMALL-FILES TABLE
| Scenario | Files | Blocks | NameNode RAM |
|---|---|---|---|
| one 1 GB file | 1 | 8 | 0.00 MB (1,350 bytes) |
| 1,000 × 1 MB | 1,000 | 1,000 | 0.29 MB |
| 1,000,000 × 1 KB | 1,000,000 | 1,000,000 | 286.10 MB |
A factor of 222,222 for the same gigabyte.
AND THE STORAGE TRADE
1 GB at replication 3 occupies 3,072 MB; under RS-6-3 erasure coding, 1,536 MB — 200% overhead against 50%, at the cost of expensive reconstruction reads.
RESULT
A 300 MB file is 128 + 128 + 44 MB, three blocks, each on three of the four DataNodes; the same file in 64 MB blocks is five. The arithmetic agrees: 260 MB is 128 + 128 + 4, and a million 1 KB files cost the NameNode 222,222 times the memory of one 1 GB file.
Simulate NameNode and DataNode failure, and observe fault tolerance and recovery.
Kill a DataNode and read the file anyway; wait for the NameNode to declare it dead and re-replicate; kill the NameNode and watch it recover from its image and edits; then work out which failures lose data.
On the cluster, 05_fault_tolerance.sh:
The Python check, 05_fault_tolerance.py:
WHICH FAILURES LOSE DATA
| Failure | Blocks live | Blocks lost |
|---|---|---|
| 1 DataNode (n1) | 8 | 0 |
| 2 DataNodes (n1, n3) | 8 | 0 |
| 3 DataNodes (n1, n3, n5) | 8 | 0 |
| 3 DataNodes (n0, n1, n3) | 6 | 2 |
| a whole rack (r1) | 8 | 0 |
| both racks | 0 | 8 |
Rows 3 and 4 are the point. n1, n3, n5 are rack 1, so rows 3 and 5
are the same failure written two ways — and both are survivable. Three
failures only hurt when they straddle the racks.
On the cluster, 05_fault_tolerance.sh:
# Experiment 5 -- simulate NameNode/DataNode failure and observe fault tolerance and recovery
# Run it: bash 05_fault_tolerance.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
# The runnable half is 05_fault_tolerance.py, which models which blocks survive which failures
# Step 1: Have the 300 MB file
hdfs dfs -test -e /user/student/big/big.bin || {
dd if=/dev/urandom of=big.bin bs=1M count=300 status=none
hdfs dfs -mkdir -p /user/student/big && hdfs dfs -put big.bin /user/student/big/; }
# Step 2: Kill a DataNode, and read the file
hdfs dfsadmin -report 2>/dev/null | grep -E "^Name:|Live datanodes"
hdfs fsck /user/student/big/big.bin -files -blocks -locations > before.txt 2>/dev/null
grep -o "DatanodeInfoWithStorage" before.txt | wc -l # 3 blocks x 3 replicas = 9
# kill one DataNode. jps shows four here -- this cluster runs four on one
# machine, as a multi-node cluster runs one per host -- so pick it by its pid file
jps | grep -c DataNode
kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-datanode.pid)
# [Corrected: `kill -9 <datanode_pid>` was a placeholder. A DataNode writes its
# pid to $HADOOP_PID_DIR (by default /tmp) as hadoop-<user>-datanode.pid.]
# the file is STILL READABLE, immediately -- other replicas serve it
hdfs dfs -cat /user/student/big/big.bin 2>client.log | wc -c
# every byte. A client tries a block's replicas in turn: if it tries the dead
# node first it logs a warning (kept in client.log here) and reads another
# replica. Which it tries first differs from read to read.
# the NameNode does not react at once: dfs.heartbeat.interval (3 s) and
# dfs.namenode.heartbeat.recheck-interval (5 min) give
# 10 * 3 + 2 * 300 = 630 seconds before the node is declared DEAD
# This cluster sets the recheck interval to 15 s, so 10 * 3 + 2 * 15 = 60 s:
# Step 3: Wait for it to be declared dead
sleep 75
hdfs dfsadmin -report -dead 2>/dev/null | grep -E "Dead datanodes"
hdfs fsck / 2>/dev/null | grep -E "Under-replicated|Missing blocks:"
# under-replicated blocks appear, then disappear as HDFS re-replicates --
# on a cluster this small, within seconds of the node being declared dead,
# so by now the copies are made. fsck has where every replica now lives:
hdfs fsck /user/student/big/big.bin -files -blocks -locations 2>/dev/null \
| grep -oE "DatanodeInfoWithStorage\[[0-9.:]+" | sort | uniq -c
# 3 blocks x 3 replicas on the three nodes still alive: each has all three.
# Whichever blocks the dead node held were copied again.
# [Corrected: this counted the NameNode's "to replicate blk_" log lines, as
# Hadoop 2 logged each order. Hadoop 3 logs them only at DEBUG, so the count
# is 0 however many blocks were copied; where the replicas are is the evidence.]
# [Corrected: this waited 660 s, for the default 630 s. The wait follows the
# setting; dfsadmin -report -dead lists only the dead.]
# Step 4: Bring it back
# bring it back
$HADOOP_HOME/bin/hdfs --daemon start datanode
sleep 15
hdfs fsck / 2>/dev/null | grep -E "Over-replicated" # briefly over-replicated, then trimmed
# Step 5: Kill the NameNode, and restart it
jps | grep -c NameNode # NameNode and SecondaryNameNode
kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-namenode.pid)
hdfs dfs -ls / 2>&1 | grep -m 1 -o "Call From .* failed on connection exception"
# FAILS. The cluster is unusable. Nothing was lost,
# but nothing is reachable either.
$HADOOP_HOME/bin/hdfs --daemon start namenode
sleep 5
hdfs dfsadmin -safemode get # ON -- it is collecting block reports
hdfs dfsadmin -safemode wait
# safe mode leaves once dfs.namenode.safemode.threshold-pct (0.999) of
# blocks have reported. On a large cluster this takes MINUTES, and it is
# why HA exists.
hdfs dfs -cat /user/student/big/big.bin | wc -c # every byte, still
# Step 6: Read what recovery reads
ls $(hdfs getconf -confKey dfs.namenode.name.dir | sed 's|^file://||')/current/ | sed -E 's/[0-9]{19}/N/g' | sort -u
# fsimage_N the namespace at a checkpoint
# edits_N-N, edits_inprogress_N every change since
# VERSION clusterID -- must match the DataNodes'
# [Corrected: this listed /usr/local/hadoop_store/hdfs/namenode/current, the
# directory experiment 1 configured on its machine. hdfs getconf asks the
# cluster where its own is; the transaction numbers are shown as N.]
# The BLOCK MAP is in NONE of these. It is rebuilt from block reports.
The Python check, 05_fault_tolerance.py:
"""Experiment 5 -- simulate NameNode/DataNode failure and observe fault
tolerance and recovery.
`05_fault_tolerance.sh` carries the commands. What runs here is the model:
which blocks survive which failures, and why the NameNode is the one failure
that is different in kind.
"""
import itertools
from blocks import blocks_for, placement
def surviving(plan, rack_of, dead_nodes=(), dead_racks=()):
"""Blocks with at least one live replica, and blocks fully lost."""
dead = set(dead_nodes) | {d for d in rack_of if rack_of[d] in dead_racks}
live, lost = [], []
for i, nodes in enumerate(plan):
(live if any(n not in dead for n in nodes) else lost).append(i)
return live, lost
def main():
print(" Experiment 5 -- fault tolerance and recovery")
# Step 1: Fail nodes and racks
n, _ = blocks_for(1024 * 1024 * 1024)
plan, rack_of = placement(n, 3, datanodes=6, racks=2)
print(f"\n a 1 GB file: {n} blocks, replication 3, 6 nodes, 2 racks")
print(f"\n {'failure':<28}{'blocks live':>12}{'blocks lost':>12} verdict")
scenarios = [
("1 DataNode (n1)", dict(dead_nodes=[1])),
("2 DataNodes (n1, n3)", dict(dead_nodes=[1, 3])),
("3 DataNodes (n1, n3, n5)", dict(dead_nodes=[1, 3, 5])),
("3 DataNodes (n0, n1, n3)", dict(dead_nodes=[0, 1, 3])),
("a whole rack (r1)", dict(dead_racks=[1])),
("both racks", dict(dead_racks=[0, 1])),
]
for label, kw in scenarios:
live, lost = surviving(plan, rack_of, **kw)
verdict = "no data loss" if not lost else f"DATA LOSS on {len(lost)}"
print(f" {label:<28}{len(live):>12}{len(lost):>12} {verdict}")
live, lost = surviving(plan, rack_of, dead_racks=[1])
assert lost == [], "losing one whole rack must not lose data"
live, lost = surviving(plan, rack_of, dead_nodes=[1, 3, 5])
assert lost == [], "n1, n3, n5 IS rack 1 -- the same failure, renamed"
print(""" losing an ENTIRE RACK loses nothing, because every block
keeps one replica on the other rack. That is precisely what
the placement policy bought, and it is the answer to 'why
rack awareness?'
Note rows 3 and 4: n1, n3, n5 ARE rack 1, so those are the
same failure written two ways -- and both are survivable.
Three failures only hurt when they straddle the racks, as
(n0, n1, n3) does""")
# Step 2: Find the worst case
print("\n the worst case, by brute force -- how many DataNode failures")
print(" can this layout survive with certainty?")
worst = None
for k in range(1, 7):
bad = [combo for combo in itertools.combinations(range(6), k)
if surviving(plan, rack_of, dead_nodes=combo)[1]]
total = len(list(itertools.combinations(range(6), k)))
print(f" {k} node(s) down: {len(bad):>3} of {total:>3} "
f"combinations lose data")
if bad and worst is None:
worst = k
assert worst == 3, "replication 3 tolerates ANY 2 failures, not any 3"
print(""" ANY TWO failures are survivable; some threes are not.
Replication factor R tolerates R-1 arbitrary failures --
and note that most 3-node combinations are still fine, so
'replication 3 fails at 3 nodes' is only true of the worst
case, which is the honest way to state it""")
# Step 3: Re-replicate
print("\n re-replication after a DataNode is declared dead:")
print(" 1. DataNode misses heartbeats (default: 3 sec interval)")
print(" 2. NameNode waits 10 * 3 sec + 2 * 5 min = 10 min 30 sec")
print(" 3. its blocks are now UNDER-REPLICATED (2 of 3)")
print(" 4. NameNode schedules copies from surviving replicas")
print(" 5. replication returns to 3; no client ever saw an error")
stale = 10 * 3 + 2 * 5 * 60
assert stale == 630
print(f""" the {stale}-second default is deliberately LONG. A node
that reboots in five minutes should not trigger a cluster-wide
copy storm, so HDFS trades a longer window of reduced
redundancy for far less needless network traffic""")
# Step 4: Lose the NameNode
print("\n the NameNode is a different kind of failure:")
print(f" {'component':<22}{'holds':<34}{'lost on crash?'}")
for comp, holds, lost_ in (
("fsimage (on disk)", "the namespace at a checkpoint", "no"),
("edit log (on disk)", "changes since the checkpoint", "no"),
("block map (in RAM)", "block -> DataNode locations", "YES"),
):
print(f" {comp:<22}{holds:<34}{lost_}")
print(""" the block MAP is never persisted -- it is rebuilt from
DataNode block reports at startup, which is why a large
NameNode takes minutes to leave safe mode. The namespace
survives; the locations are reconstructed""")
# Step 5: Compare the three answers
print("\n the three answers to NameNode failure, in historical order:")
print(f" {'mechanism':<26}{'recovers':<16}{'automatic?'}")
for m, r, a in (("Secondary NameNode", "checkpoint only", "no -- NOT a standby"),
("NameNode HA (2 NNs)", "full", "yes, via ZooKeeper"),
("HDFS Federation", "n/a -- scales namespace", "n/a")):
print(f" {m:<26}{r:<16}{a}")
print(""" the Secondary NameNode is the most misleadingly named
component in Hadoop: it merges fsimage with the edit log so
restarts stay fast, and it CANNOT take over. HA needs two
NameNodes, a shared edit log (QJM) and ZooKeeper for failover
-- which is exactly why experiment 16 exists""")
if __name__ == "__main__":
main()
On the cluster, 05_fault_tolerance.sh:
OUTPUT
$ hdfs dfs -test -e /user/student/big/big.bin || {
dd if=/dev/urandom of=big.bin bs=1M count=300 status=none
hdfs dfs -mkdir -p /user/student/big && hdfs dfs -put big.bin /user/student/big/; }
$ hdfs dfsadmin -report 2>/dev/null | grep -E "^Name:|Live datanodes"
Live datanodes (4):
Name: 127.0.0.1:9866 (localhost)
Name: 127.0.0.1:9886 (localhost)
Name: 127.0.0.1:9896 (localhost)
Name: 127.0.0.1:9906 (localhost)
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations > before.txt 2>/dev/null
$ grep -o "DatanodeInfoWithStorage" before.txt | wc -l
9
$ jps | grep -c DataNode
4
$ kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-datanode.pid)
$ hdfs dfs -cat /user/student/big/big.bin 2>client.log | wc -c
314572800
$ sleep 75
$ hdfs dfsadmin -report -dead 2>/dev/null | grep -E "Dead datanodes"
Dead datanodes (1):
$ hdfs fsck / 2>/dev/null | grep -E "Under-replicated|Missing blocks:"
Under-replicated blocks: 0 (0.0 %)
Missing blocks: 0
$ hdfs fsck /user/student/big/big.bin -files -blocks -locations 2>/dev/null \
| grep -oE "DatanodeInfoWithStorage\[[0-9.:]+" | sort | uniq -c
3 DatanodeInfoWithStorage[127.0.0.1:9886
3 DatanodeInfoWithStorage[127.0.0.1:9896
3 DatanodeInfoWithStorage[127.0.0.1:9906
$ $HADOOP_HOME/bin/hdfs --daemon start datanode
$ sleep 15
$ hdfs fsck / 2>/dev/null | grep -E "Over-replicated"
Over-replicated blocks: 0 (0.0 %)
$ jps | grep -c NameNode
2
$ kill -9 $(cat ${HADOOP_PID_DIR:-/tmp}/hadoop-$USER-namenode.pid)
$ hdfs dfs -ls / 2>&1 | grep -m 1 -o "Call From .* failed on connection exception"
Call From vm/127.0.0.1 to localhost:9000 failed on connection exception
$ $HADOOP_HOME/bin/hdfs --daemon start namenode
$ sleep 5
$ hdfs dfsadmin -safemode get
Safe mode is ON
$ hdfs dfsadmin -safemode wait
Safe mode is OFF
$ hdfs dfs -cat /user/student/big/big.bin | wc -c
314572800
$ ls $(hdfs getconf -confKey dfs.namenode.name.dir | sed 's|^file://||')/current/ | sed -E 's/[0-9]{19}/N/g' | sort -u
VERSION
edits_N-N
edits_inprogress_N
fsimage_N
fsimage_N.md5
seen_txid
The Python check, 05_fault_tolerance.py:
OUTPUT
Experiment 5 -- fault tolerance and recovery
a 1 GB file: 8 blocks, replication 3, 6 nodes, 2 racks
failure blocks live blocks lost verdict
1 DataNode (n1) 8 0 no data loss
2 DataNodes (n1, n3) 8 0 no data loss
3 DataNodes (n1, n3, n5) 8 0 no data loss
3 DataNodes (n0, n1, n3) 6 2 DATA LOSS on 2
a whole rack (r1) 8 0 no data loss
both racks 0 8 DATA LOSS on 8
losing an ENTIRE RACK loses nothing, because every block
keeps one replica on the other rack. That is precisely what
the placement policy bought, and it is the answer to 'why
rack awareness?'
Note rows 3 and 4: n1, n3, n5 ARE rack 1, so those are the
same failure written two ways -- and both are survivable.
Three failures only hurt when they straddle the racks, as
(n0, n1, n3) does
the worst case, by brute force -- how many DataNode failures
can this layout survive with certainty?
1 node(s) down: 0 of 6 combinations lose data
2 node(s) down: 0 of 15 combinations lose data
3 node(s) down: 6 of 20 combinations lose data
4 node(s) down: 12 of 15 combinations lose data
5 node(s) down: 6 of 6 combinations lose data
6 node(s) down: 1 of 1 combinations lose data
ANY TWO failures are survivable; some threes are not.
Replication factor R tolerates R-1 arbitrary failures --
and note that most 3-node combinations are still fine, so
'replication 3 fails at 3 nodes' is only true of the worst
case, which is the honest way to state it
re-replication after a DataNode is declared dead:
1. DataNode misses heartbeats (default: 3 sec interval)
2. NameNode waits 10 * 3 sec + 2 * 5 min = 10 min 30 sec
3. its blocks are now UNDER-REPLICATED (2 of 3)
4. NameNode schedules copies from surviving replicas
5. replication returns to 3; no client ever saw an error
the 630-second default is deliberately LONG. A node
that reboots in five minutes should not trigger a cluster-wide
copy storm, so HDFS trades a longer window of reduced
redundancy for far less needless network traffic
the NameNode is a different kind of failure:
component holds lost on crash?
fsimage (on disk) the namespace at a checkpoint no
edit log (on disk) changes since the checkpoint no
block map (in RAM) block -> DataNode locations YES
the block MAP is never persisted -- it is rebuilt from
DataNode block reports at startup, which is why a large
NameNode takes minutes to leave safe mode. The namespace
survives; the locations are reconstructed
the three answers to NameNode failure, in historical order:
mechanism recovers automatic?
Secondary NameNode checkpoint only no -- NOT a standby
NameNode HA (2 NNs) full yes, via ZooKeeper
HDFS Federation n/a -- scales namespacen/a
the Secondary NameNode is the most misleadingly named
component in Hadoop: it merges fsimage with the edit log so
restarts stay fast, and it CANNOT take over. HA needs two
NameNodes, a shared edit log (QJM) and ZooKeeper for failover
-- which is exactly why experiment 16 exists
THE HONEST VERSION, BY BRUTE FORCE
| Nodes down | Combinations losing data |
|---|---|
| 1 | 0 of 6 |
| 2 | 0 of 15 |
| 3 | 6 of 20 |
| 4 | 12 of 15 |
| 5 | 6 of 6 |
Any two failures are survivable; 14 of the 20 three-node combinations are still fine. "Replication 3 fails at 3 nodes" is the worst case, not the rule, and stating it that way is the honest answer.
THE 630-SECOND DELAY
10 × 3 s + 2 × 5 min = 630 s before a DataNode is declared dead. The delay
is deliberate — a node that reboots in five minutes should not trigger a
cluster-wide copy storm. The lab's cluster sets the recheck interval to 15 s, so
10 × 3 s + 2 × 15 s = 60 s, and the script waits 75 s rather than eleven minutes; the
formula is the same.
THE NAMENODE'S BLOCK MAP IS NEVER PERSISTED
| Component | Holds | Lost on crash? |
|---|---|---|
fsimage (disk) |
the namespace at a checkpoint | no |
edits (disk) |
changes since | no |
| block map (RAM) | block → DataNode locations | YES |
Rebuilt from block reports at startup, which is why a large NameNode takes
minutes to leave safe mode — the safemode get above caught it there. And the Secondary
NameNode is a checkpointer, not a standby — the most misleadingly named component in Hadoop.
RESULT
With a DataNode killed the 300 MB file was still read in full; the node was declared dead and its blocks re-replicated to the three still alive; with the NameNode killed nothing was reachable, and after its restart and safe mode every byte was there. In the model, any two failures are survivable and 6 of the 20 three-node failures lose data.
Configure YARN and run sample applications, observing the roles of the ResourceManager and the NodeManager.
Set up two Capacity Scheduler queues, run the sample applications into each, observe the nodes and queues, and kill a running job; then compare FIFO, fair and capacity scheduling on one workload.
On the cluster, 06_yarn.sh:
The Python check, 06_yarn_scheduling.py:
THE SCHEDULERS, ON ONE WORKLOAD
The workload: an 8-container cluster; big_etl needs all 8 for 10 s;
small_q1 and small_q2 need 1 container for 2 s; medium needs 4 for 5 s.
| Job | FIFO | Fair | Capacity (75/25) |
|---|---|---|---|
big_etl |
10 | 14 | 20 |
small_q1 |
11 | 1 | 2 |
small_q2 |
12 | 1 | 3 |
medium |
16 | 5 | 22 |
| total turnaround | 49 | 21 | — |
On the cluster, 06_yarn.sh:
# Experiment 6 -- configure YARN, run sample applications, observe ResourceManager and NodeManager roles
#
# Run it: bash 06_yarn.sh, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 06_yarn_scheduling.py, which runs FIFO, Fair and Capacity on one workload
#
# Step 1: Read the configuration that matters
# yarn.nodemanager.resource.memory-mb 8192 RAM this node offers
# yarn.nodemanager.resource.cpu-vcores 8 cores this node offers
# yarn.scheduler.minimum-allocation-mb 1024 container granularity
# yarn.scheduler.maximum-allocation-mb 8192 biggest single container
# yarn.resourcemanager.scheduler.class
# org.apache.hadoop.yarn.server.resourcemanager.scheduler.
# capacity.CapacityScheduler
# yarn.nodemanager.aux-services mapreduce_shuffle
# ^ without this the shuffle has no server and every job hangs at 33%
# Step 2: Set up the queues
# yarn.scheduler.capacity.root.queues production,adhoc
# yarn.scheduler.capacity.root.production.capacity 75
# yarn.scheduler.capacity.root.adhoc.capacity 25
# yarn.scheduler.capacity.root.adhoc.maximum-capacity 50 <- ELASTICITY
yarn rmadmin -refreshQueues # queues reload WITHOUT restarting the RM
yarn queue -status production 2>/dev/null | grep -E "Queue Name|Capacity"
yarn queue -status adhoc 2>/dev/null | grep -E "Queue Name|Capacity"
# Step 3: Run the sample applications
EX=$(ls $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar)
yarn jar $EX pi 4 1000 2>&1 | grep -E "Estimated value|Job Finished"
yarn jar $EX teragen 100000 /user/student/terasort-in 2>&1 | grep -E "completed successfully"
yarn jar $EX terasort /user/student/terasort-in /user/student/terasort-out 2>&1 | grep -E "completed successfully"
yarn jar $EX teravalidate /user/student/terasort-out /user/student/terasort-check 2>&1 | grep -E "completed successfully"
hdfs dfs -cat /user/student/terasort-check/part-r-00000
# no "misorder" line: the output is in order. The checksum is of the rows.
# [Corrected: teragen wrote 10,000,000 rows -- a gigabyte, three on disk with
# replication 3 -- which is slow on one machine and proves nothing more than
# 100,000 rows (10 MB) do. teravalidate is added: it is what says the sort worked.]
# Step 4: Submit to a named queue
# submit into a named queue and watch where it lands
yarn jar $EX pi -Dmapreduce.job.queuename=adhoc 4 1000 2>&1 | grep -E "Estimated value"
yarn application -list -appStates FINISHED 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $5 " | " $7}' | sort
# [Corrected: this was `yarn jar ...examples-*.jar`, with three dots typed
# where the path goes; the path is set once, in EX, above.]
# Step 5: Observe the nodes and queues
# a job's client returns when the job reports success, and its last container is
# released a moment later; wait for that, so what follows is the idle cluster
until yarn node -list 2>/dev/null | grep -qE "RUNNING.*[[:space:]]0$"; do sleep 1; done
yarn node -list -all 2>/dev/null | tail -n +2 # every NodeManager, its state and containers
yarn queue -status adhoc 2>/dev/null | grep -E "Current Capacity|Maximum Capacity"
# Step 6: Kill a job
# a job to kill: start one that runs a while, and kill it by its id
yarn jar $EX pi 4 1000000000 > /dev/null 2>&1 &
sleep 15
APP=$(yarn application -list 2>/dev/null | awk '/^application_/{print $1}' | head -1)
yarn application -list 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $6}'
yarn application -kill $APP 2>&1 | grep -E "Killed application|Killing application"
wait
yarn application -status $APP 2>/dev/null | grep -E "Final-State"
# [Corrected: this killed application_1699999999999_0002, an id from someone
# else's cluster; the id is taken from the list.]
#
# yarn top # like top(1), for the cluster
# -- a full-screen display that runs until you press q, so it is not run here.
# http://localhost:8088/cluster/scheduler the queue tree, live
#
# What to look for: submit a big job and a small one into different queues,
# and watch the adhoc job start IMMEDIATELY even though production is full.
# That guarantee is what a queue is, and it is the whole answer to
# "what does the Capacity Scheduler do".
The Python check, 06_yarn_scheduling.py:
"""Experiment 6 -- configure YARN, run sample applications, and observe the
ResourceManager and NodeManager roles.
`06_yarn.sh` carries the real commands. What runs here is the SCHEDULER, which
is the part of YARN that actually decides anything -- and the part where the
three policies give visibly different answers on the same queue.
"""
# (name, containers needed, seconds per container-slot, submitted at)
JOBS = [
("big_etl", 8, 10, 0),
("small_q1", 1, 2, 1),
("small_q2", 1, 2, 2),
("medium", 4, 5, 3),
]
CLUSTER = 8 # containers available cluster-wide
def fifo(jobs, capacity):
"""First in, first out. One job owns the cluster until it finishes."""
now, done = 0, {}
for name, need, secs, submitted in sorted(jobs, key=lambda j: j[3]):
start = max(now, submitted)
waves = -(-need // capacity) # ceil
now = start + waves * secs
done[name] = (start, now, now - submitted)
return done
def fair(jobs, capacity):
"""Fair scheduler: every RUNNING job gets an equal share of containers.
Simulated one second at a time, which is crude but exactly right for
showing the property that matters -- a one-container job does not wait
behind an eight-container job.
"""
remaining = {j[0]: j[1] * j[2] for j in jobs} # container-seconds of work
submitted = {j[0]: j[3] for j in jobs}
done, t = {}, 0
while any(v > 0 for v in remaining.values()):
active = [n for n, v in remaining.items() if v > 0 and submitted[n] <= t]
if not active:
t += 1
continue
share = capacity / len(active)
for n in active:
remaining[n] -= share
if remaining[n] <= 0 and n not in done:
done[n] = (submitted[n], t + 1, t + 1 - submitted[n])
t += 1
return done
def capacity_sched(jobs, capacity, queues):
"""Capacity scheduler: queues get guaranteed percentages of the cluster.
A job cannot exceed its queue's share even when the cluster is idle,
unless elasticity is enabled -- which is the whole difference between
'capacity' and 'fair'.
"""
done = {}
for qname, pct, members in queues:
slots = max(1, int(capacity * pct / 100))
qjobs = [j for j in jobs if j[0] in members]
sub = fifo(qjobs, slots)
for k, v in sub.items():
done[k] = v
return done
def main():
print(" Experiment 6 -- YARN scheduling")
# Step 1: Set out the workload
print("\n the workload, on an 8-container cluster:")
print(f" {'job':<10}{'containers':>11}{'sec/wave':>10}{'submitted':>11}")
for name, need, secs, sub in JOBS:
print(f" {name:<10}{need:>11}{secs:>10}{sub:>11}")
f = fifo(JOBS, CLUSTER)
# Step 2: Schedule first in, first out
print("\n FIFO scheduler:")
print(f" {'job':<10}{'start':>7}{'finish':>8}{'turnaround':>12}")
for name, _, _, _ in JOBS:
s, e, t = f[name]
print(f" {name:<10}{s:>7}{e:>8}{t:>12}")
fifo_small = f["small_q1"][2]
print(f""" small_q1 needs ONE container for TWO seconds and waits
{fifo_small} seconds, because big_etl took the whole cluster first.
That is head-of-line blocking, and it is why nobody runs
FIFO on a shared cluster""")
fr = fair(JOBS, CLUSTER)
# Step 3: Schedule fairly
print("\n Fair scheduler:")
print(f" {'job':<10}{'start':>7}{'finish':>8}{'turnaround':>12}")
for name, _, _, _ in JOBS:
s, e, t = fr[name]
print(f" {name:<10}{s:>7}{e:>8}{t:>12}")
fair_small = fr["small_q1"][2]
assert fair_small < fifo_small
print(f""" small_q1 now finishes in {fair_small}s instead of {fifo_small}s.
Fair sharing did not make the cluster faster -- big_etl
finished LATER ({f['big_etl'][1]} -> {fr['big_etl'][1]}) -- it moved latency from
the small job to the big one, which is almost always the
trade you want on an interactive cluster""")
total_fifo = sum(v[2] for v in f.values())
total_fair = sum(v[2] for v in fr.values())
work = sum(need * secs for _, need, secs, _ in JOBS)
print(f"\n total container-seconds of WORK: {work} either way")
print(f" total turnaround: FIFO {total_fifo} Fair {total_fair}")
assert total_fair < total_fifo
print(f""" the work is identical -- {work} container-seconds, which on
8 containers cannot finish before second {-(-work // CLUSTER)}. What changed is
WAITING: FIFO made three jobs queue behind one, so total
turnaround fell from {total_fifo} to {total_fair} without the cluster doing
anything faster.
Scheduling decides WHO waits. It cannot create throughput,
but idle-while-queued is real waste and fair sharing removes
it""")
cap = capacity_sched(JOBS, CLUSTER, [
("production", 75, {"big_etl", "medium"}),
("adhoc", 25, {"small_q1", "small_q2"}),
])
# Step 4: Schedule by capacity
print("\n Capacity scheduler -- production 75%, adhoc 25%:")
print(f" {'job':<10}{'queue':<12}{'start':>7}{'finish':>8}{'turnaround':>12}")
for name, q in (("big_etl", "production"), ("medium", "production"),
("small_q1", "adhoc"), ("small_q2", "adhoc")):
s, e, t = cap[name]
print(f" {name:<10}{q:<12}{s:>7}{e:>8}{t:>12}")
print(""" the adhoc queue holds 2 containers whatever else is
running, so a short query has a GUARANTEE rather than a
hope. The cost: those 2 containers sit idle when adhoc is
empty, unless queue elasticity is turned on""")
# Step 5: Name who does what
print("\n who does what in YARN:")
print(f" {'component':<22}{'one per':<14}{'responsibility'}")
for c, per, resp in (
("ResourceManager", "cluster", "global scheduling; hands out containers"),
("NodeManager", "node", "launches and monitors containers, reports health"),
("ApplicationMaster", "JOB", "negotiates containers, retries failed tasks"),
("Container", "task", "a bounded slice of CPU and RAM on one node")):
print(f" {c:<22}{per:<14}{resp}")
print(""" ONE ApplicationMaster PER JOB is the change that defined
YARN. In Hadoop 1 the JobTracker did both scheduling and job
management for every job, so it was the bottleneck AND the
single point of failure. Splitting them is why YARN can run
Spark, Tez and Flink and not only MapReduce""")
if __name__ == "__main__":
main()
On the cluster, 06_yarn.sh:
OUTPUT
$ yarn rmadmin -refreshQueues
2026-10-04 23:12:43,788 INFO client.DefaultNoHARMFailoverProxyProvider: Connecting to ResourceManager at /0.0.0.0:8033
$ yarn queue -status production 2>/dev/null | grep -E "Queue Name|Capacity"
Queue Name : production
Capacity : 75.00%
Current Capacity : .00%
Maximum Capacity : 100.00%
$ yarn queue -status adhoc 2>/dev/null | grep -E "Queue Name|Capacity"
Queue Name : adhoc
Capacity : 25.00%
Current Capacity : .00%
Maximum Capacity : 50.00%
$ EX=$(ls $HADOOP_HOME/share/hadoop/mapreduce/hadoop-mapreduce-examples-*.jar)
$ yarn jar $EX pi 4 1000 2>&1 | grep -E "Estimated value|Job Finished"
Job Finished in 24.944 seconds
Estimated value of Pi is 3.14000000000000000000
$ yarn jar $EX teragen 100000 /user/student/terasort-in 2>&1 | grep -E "completed successfully"
2026-10-04 23:13:29,432 INFO mapreduce.Job: Job job_1791155559313_0002 completed successfully
$ yarn jar $EX terasort /user/student/terasort-in /user/student/terasort-out 2>&1 | grep -E "completed successfully"
2026-10-04 23:13:51,600 INFO mapreduce.Job: Job job_1791155559313_0003 completed successfully
$ yarn jar $EX teravalidate /user/student/terasort-out /user/student/terasort-check 2>&1 | grep -E "completed successfully"
2026-10-04 23:14:11,151 INFO mapreduce.Job: Job job_1791155559313_0004 completed successfully
$ hdfs dfs -cat /user/student/terasort-check/part-r-00000
checksum c327a1c42c28
$ yarn jar $EX pi -Dmapreduce.job.queuename=adhoc 4 1000 2>&1 | grep -E "Estimated value"
Estimated value of Pi is 3.14000000000000000000
$ yarn application -list -appStates FINISHED 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $5 " | " $7}' | sort
TeraGen | production | SUCCEEDED
TeraSort | production | SUCCEEDED
TeraValidate | production | SUCCEEDED
QuasiMonteCarlo | adhoc | SUCCEEDED
QuasiMonteCarlo | production | SUCCEEDED
$ until yarn node -list 2>/dev/null | grep -qE "RUNNING.*[[:space:]]0$"; do sleep 1; done
$ yarn node -list -all 2>/dev/null | tail -n +2
Node-Id Node-State Node-Http-Address Number-of-Running-Containers
localhost:36875 RUNNING localhost:8042 0
$ yarn queue -status adhoc 2>/dev/null | grep -E "Current Capacity|Maximum Capacity"
Current Capacity : .00%
Maximum Capacity : 50.00%
$ yarn jar $EX pi 4 1000000000 > /dev/null 2>&1 &
$ sleep 15
$ APP=$(yarn application -list 2>/dev/null | awk '/^application_/{print $1}' | head -1)
$ yarn application -list 2>/dev/null | awk -F'\t' 'NR>2{print $2 " | " $6}'
QuasiMonteCarlo | RUNNING
$ yarn application -kill $APP 2>&1 | grep -E "Killed application|Killing application"
Killing application application_1791155559313_0006
2026-10-04 23:15:07,634 INFO impl.YarnClientImpl: Killed application application_1791155559313_0006
$ wait
$ yarn application -status $APP 2>/dev/null | grep -E "Final-State"
Final-State : KILLED
The Python check, 06_yarn_scheduling.py:
OUTPUT
Experiment 6 -- YARN scheduling
the workload, on an 8-container cluster:
job containers sec/wave submitted
big_etl 8 10 0
small_q1 1 2 1
small_q2 1 2 2
medium 4 5 3
FIFO scheduler:
job start finish turnaround
big_etl 0 10 10
small_q1 10 12 11
small_q2 12 14 12
medium 14 19 16
small_q1 needs ONE container for TWO seconds and waits
11 seconds, because big_etl took the whole cluster first.
That is head-of-line blocking, and it is why nobody runs
FIFO on a shared cluster
Fair scheduler:
job start finish turnaround
big_etl 0 14 14
small_q1 1 2 1
small_q2 2 3 1
medium 3 8 5
small_q1 now finishes in 1s instead of 11s.
Fair sharing did not make the cluster faster -- big_etl
finished LATER (10 -> 14) -- it moved latency from
the small job to the big one, which is almost always the
trade you want on an interactive cluster
total container-seconds of WORK: 104 either way
total turnaround: FIFO 49 Fair 21
the work is identical -- 104 container-seconds, which on
8 containers cannot finish before second 13. What changed is
WAITING: FIFO made three jobs queue behind one, so total
turnaround fell from 49 to 21 without the cluster doing
anything faster.
Scheduling decides WHO waits. It cannot create throughput,
but idle-while-queued is real waste and fair sharing removes
it
Capacity scheduler -- production 75%, adhoc 25%:
job queue start finish turnaround
big_etl production 0 20 20
medium production 20 25 22
small_q1 adhoc 1 3 2
small_q2 adhoc 3 5 3
the adhoc queue holds 2 containers whatever else is
running, so a short query has a GUARANTEE rather than a
hope. The cost: those 2 containers sit idle when adhoc is
empty, unless queue elasticity is turned on
who does what in YARN:
component one per responsibility
ResourceManager cluster global scheduling; hands out containers
NodeManager node launches and monitors containers, reports health
ApplicationMaster JOB negotiates containers, retries failed tasks
Container task a bounded slice of CPU and RAM on one node
ONE ApplicationMaster PER JOB is the change that defined
YARN. In Hadoop 1 the JobTracker did both scheduling and job
management for every job, so it was the bottleneck AND the
single point of failure. Splitting them is why YARN can run
Spark, Tez and Flink and not only MapReduce
WHAT THE NUMBERS SAY, STATED CAREFULLY
small_q1 waits 11 s under FIFO for 2 s of work. Head-of-line blocking.Fair sharing did not make the cluster faster. big_etl finished later,
10 → 14. Latency moved from the small jobs to the big one.
The work is identical: 104 container-seconds either way, which on 8 containers cannot finish before second 13.
Total turnaround still halved, 49 → 21, because FIFO left three jobs idle in a queue. Scheduling cannot create throughput, but idle-while-queued is real waste.
THE CAPACITY SCHEDULER'S GUARANTEE
The adhoc queue holds 25% of the cluster whatever else is running, so a short
query has a guarantee rather than a hope — the queues the script set up report 75% and 25%,
with adhoc allowed to grow to 50%. The cost: that share sits idle when adhoc is empty, unless
maximum-capacity is raised to allow elasticity — which is exactly the difference between
"capacity" and "fair".
One ApplicationMaster per job is the change that defined YARN. Hadoop 1's JobTracker did both scheduling and per-job management, so it was the bottleneck and the single point of failure, and it could run only MapReduce.
RESULT
Pi, TeraGen, TeraSort and TeraValidate ran in the production queue and Pi again in adhoc, each SUCCEEDED; a long Pi job was killed by its id and finished KILLED. In the model, fair sharing halves total turnaround, 49 s to 21 s, on the same 104 container-seconds of work.
Write a simple MapReduce program for word count.
Count the words in six documents with a Java MapReduce job, with a combiner; then make the shuffle visible with a MapReduce engine in Python, and see when a combiner is safe.
On the cluster, WordCount.java:
The Python check, 07_wordcount.py:
THE SHUFFLE, MADE VISIBLE
mapreduce.py, which 07_wordcount.py imports, is a MapReduce engine in forty lines, written
out in full. The point is that it makes the shuffle visible, and the shuffle is the part
students never see and the part that costs the money.
| Phase | Records |
|---|---|
| map output | 48 |
| shuffled | 48 |
| reduce output | 26 |
Top words: the 5, big 4, data 4, dog 4, quick 3, fox 3.
The counts sum back to 48 — reduce is a regrouping, and if your totals do
not reconcile, your reducer is not associative.
On the cluster, WordCount.java:
// Experiment 7 -- word count in MapReduce
//
// Run it: the build-and-run lines below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 07_wordcount.py, which runs the same map and reduce
// functions through an explicit engine and asserts every count
//
// Build and run:
// javac -classpath $(hadoop classpath) -d classes WordCount.java
// jar -cvf WordCount.jar -C classes/ .
// hadoop jar WordCount.jar WordCount /user/student/docs /user/student/out
//
import java.io.IOException;
import java.util.StringTokenizer;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
public class WordCount {
// Step 1: Map each word to 1
public static class TokenizerMapper
extends Mapper<LongWritable, Text, Text, IntWritable> {
// Reused across every call. Allocating a new Text per word would
// create one object per word in the corpus, and the GC pause is the
// job. This is the single most important idiom in MapReduce Java.
private final static IntWritable ONE = new IntWritable(1);
private final Text word = new Text();
@Override
public void map(LongWritable key, Text value, Context context)
throws IOException, InterruptedException {
// key is the BYTE OFFSET of the line, not a line number.
StringTokenizer itr = new StringTokenizer(value.toString());
while (itr.hasMoreTokens()) {
word.set(itr.nextToken().toLowerCase());
context.write(word, ONE);
}
}
}
// Step 2: Reduce by summing
public static class IntSumReducer
extends Reducer<Text, IntWritable, Text, IntWritable> {
private final IntWritable result = new IntWritable();
@Override
public void reduce(Text key, Iterable<IntWritable> values, Context context)
throws IOException, InterruptedException {
int sum = 0;
for (IntWritable val : values) {
sum += val.get();
}
// The Iterable is streamed from disk and can be walked ONCE.
// Calling values.iterator() a second time yields nothing -- the
// classic bug when someone tries to compute a mean and a count.
result.set(sum);
context.write(key, result);
}
}
// Step 3: Configure the job, with a combiner
public static void main(String[] args) throws Exception {
Configuration conf = new Configuration();
Job job = Job.getInstance(conf, "word count");
job.setJarByClass(WordCount.class);
job.setMapperClass(TokenizerMapper.class);
job.setCombinerClass(IntSumReducer.class); // safe: sum is associative
job.setReducerClass(IntSumReducer.class);
job.setOutputKeyClass(Text.class);
job.setOutputValueClass(IntWritable.class);
job.setNumReduceTasks(1);
FileInputFormat.addInputPath(job, new Path(args[0]));
FileOutputFormat.setOutputPath(job, new Path(args[1]));
// The output directory MUST NOT EXIST. Hadoop refuses to overwrite,
// which prevents a re-run from silently destroying yesterday's result.
System.exit(job.waitForCompletion(true) ? 0 : 1);
}
}
// Expected output on the six sample documents (verified in 07_wordcount.py):
// the 5, big 4, data 4, dog 4, fox 3, quick 3, ... 26 terms, 48 words
//
// The combiner is the same class as the reducer here ONLY because sum is
// associative and commutative. Setting a combiner for a mean silently
// produces wrong answers -- mean of means is not the mean -- and Hadoop will
// not warn you.
The Python check, 07_wordcount.py:
"""Experiment 7 -- a simple MapReduce program for word count.
This is the "hello world" of MapReduce, and it is worth more than it looks:
the shuffle between map and reduce is the only part of the model that costs
real money on a cluster, and word count is the smallest program that makes it
visible.
Runs through the engine in mapreduce.py, and then through REAL PYSPARK if the
Spark virtual environment is present (see tools/setup_spark.sh).
"""
from mapreduce import run
import fixtures as f
INPUT = sorted(f.DOCS.items())
def mapper(name, line):
"""(filename, line) -> (word, 1) for every word."""
for word in line.split():
yield word, 1
def reducer(word, counts):
"""(word, [1, 1, ...]) -> (word, total)."""
yield word, sum(counts)
def main():
print(" Experiment 7 -- word count in MapReduce")
print(f"\n input: {len(INPUT)} documents, "
f"{sum(len(t.split()) for _, t in INPUT)} words")
# Step 1: Map, shuffle and reduce the documents
trace = {}
result = run(INPUT, mapper, reducer, trace=trace)
counts = dict(result)
print(f"\n {'phase':<26}{'records':>9}")
print(f" {'map output':<26}{trace['map_output']:>9}")
print(f" {'shuffled across network':<26}{trace['shuffled']:>9}")
print(f" {'reduce output':<26}{len(result):>9}")
assert trace["map_output"] == 48
assert len(result) == 26
# Step 2: Read the top words
print("\n the top words:")
for w, c in sorted(counts.items(), key=lambda kv: (-kv[1], kv[0]))[:6]:
print(f" {w:<10}{c:>3}")
assert counts["the"] == 5 and counts["dog"] == 4
assert counts["big"] == 4 and counts["data"] == 4
assert sum(counts.values()) == 48
print(""" the counts sum back to 48, the map output. Nothing was
created or lost -- reduce is a REGROUPING, and if your
totals do not reconcile, your reducer is not associative""")
# Step 3: Add a combiner
ctrace = {}
combined = run(INPUT, mapper, reducer, combiner=reducer, trace=ctrace)
assert combined == result, "a combiner must not change the answer"
saved = trace["shuffled"] - ctrace["shuffled"]
pct = 100 * saved / trace["shuffled"]
print(f"\n with a combiner (the reducer, run map-side, PER TASK):")
print(f" shuffled {trace['shuffled']} -> {ctrace['shuffled']} "
f"({saved} fewer records, {pct:.2f}%)")
print(f""" same answer, {pct:.1f}% less network. And note how SMALL that
saving is: these documents are 5 to 11 words, so there is
almost nothing to merge within one split. On a 128 MB split
of real text the same combiner cuts the shuffle by orders of
magnitude. The combiner's value scales with SPLIT SIZE, which
is the point this tiny dataset makes by failing to impress""")
# Step 4: See when a combiner is not safe
print("\n when a combiner is NOT safe:")
print(f" {'reducer computes':<22}{'combiner-safe?':<16}why")
for what, safe, why in (
("sum", "yes", "associative and commutative"),
("max", "yes", "max of maxes is the max"),
("count", "yes", "if the combiner emits partial counts"),
("MEAN", "NO", "mean of means is not the mean"),
("median", "NO", "needs every value at once")):
print(f" {what:<22}{safe:<16}{why}")
# prove the mean case rather than asserting it
groups = [[1, 1, 1, 10], [10]]
naive = sum(sum(g) / len(g) for g in groups) / len(groups)
true = sum(sum(g) for g in groups) / sum(len(g) for g in groups)
print(f"\n mean of means = {naive:.4f}, true mean = {true:.4f}")
assert abs(naive - true) > 1
print(f""" {naive:.4f} against {true:.4f} on five numbers. To average safely,
emit (sum, count) pairs from the combiner and divide only in
the reducer -- and that is the same average-of-averages trap
Course 11 met in DAX, in a different costume""")
# Step 5: Partition the keys
print("\n 3 reduce tasks instead of 1:")
ptrace = {}
three = run(INPUT, mapper, reducer, reducers=3, trace=ptrace)
assert three == result, "the number of reducers must not change the answer"
sizes = ptrace["partition_sizes"]
print(f" partition sizes: {sizes} (total {sum(sizes)})")
print(f" largest / smallest = {max(sizes) / min(sizes):.2f}")
print(""" hash partitioning is only as balanced as the KEY
DISTRIBUTION. Natural language is Zipfian, so a real corpus
skews far worse than this -- one reducer gets 'the' and
finishes last, and the job's wall clock is that reducer.
Skew, not volume, is what usually kills a MapReduce job""")
return counts
if __name__ == "__main__":
main()
On the cluster, WordCount.java:
OUTPUT
$ hdfs dfs -mkdir -p /user/student/docs
$ hdfs dfs -put docs/*.txt /user/student/docs/
$ hdfs dfs -ls /user/student/docs | awk 'NR>1{print $NF}'
/user/student/docs/doc1.txt
/user/student/docs/doc2.txt
/user/student/docs/doc3.txt
/user/student/docs/doc4.txt
/user/student/docs/doc5.txt
/user/student/docs/doc6.txt
$ mkdir -p classes
$ javac -classpath $(hadoop classpath) -d classes WordCount.java
$ jar -cvf WordCount.jar -C classes/ .
added manifest
adding: WordCount$TokenizerMapper.class(in = 1895) (out= 799)(deflated 57%)
adding: WordCount.class(in = 1530) (out= 832)(deflated 45%)
adding: WordCount$IntSumReducer.class(in = 1739) (out= 740)(deflated 57%)
$ hadoop jar WordCount.jar WordCount /user/student/docs /user/student/out 2>&1 | grep -E "completed successfully|Map input records|Map output records|Combine input records|Combine output records|Reduce input groups|Reduce output records"
2026-10-04 23:16:12,108 INFO mapreduce.Job: Job job_1791155734078_0001 completed successfully
Map input records=6
Map output records=48
Combine input records=48
Combine output records=39
Reduce input groups=26
Reduce output records=26
$ hdfs dfs -cat /user/student/out/part-r-00000
a 2
all 1
and 2
big 4
brown 2
data 4
day 1
dog 4
for 1
fox 3
hadoop 1
is 2
jumps 1
lazy 2
machine 1
one 1
outpaces 1
over 1
processes 1
quick 3
sleeps 1
spark 1
stores 1
that 1
the 5
too 1
The Python check, 07_wordcount.py:
OUTPUT
Experiment 7 -- word count in MapReduce
input: 6 documents, 48 words
phase records
map output 48
shuffled across network 48
reduce output 26
the top words:
the 5
big 4
data 4
dog 4
fox 3
quick 3
the counts sum back to 48, the map output. Nothing was
created or lost -- reduce is a REGROUPING, and if your
totals do not reconcile, your reducer is not associative
with a combiner (the reducer, run map-side, PER TASK):
shuffled 48 -> 39 (9 fewer records, 18.75%)
same answer, 18.8% less network. And note how SMALL that
saving is: these documents are 5 to 11 words, so there is
almost nothing to merge within one split. On a 128 MB split
of real text the same combiner cuts the shuffle by orders of
magnitude. The combiner's value scales with SPLIT SIZE, which
is the point this tiny dataset makes by failing to impress
when a combiner is NOT safe:
reducer computes combiner-safe? why
sum yes associative and commutative
max yes max of maxes is the max
count yes if the combiner emits partial counts
MEAN NO mean of means is not the mean
median NO needs every value at once
mean of means = 6.6250, true mean = 4.6000
6.6250 against 4.6000 on five numbers. To average safely,
emit (sum, count) pairs from the combiner and divide only in
the reducer -- and that is the same average-of-averages trap
Course 11 met in DAX, in a different costume
3 reduce tasks instead of 1:
partition sizes: [20, 17, 11] (total 48)
largest / smallest = 1.82
hash partitioning is only as balanced as the KEY
DISTRIBUTION. Natural language is Zipfian, so a real corpus
skews far worse than this -- one reducer gets 'the' and
finishes last, and the job's wall clock is that reducer.
Skew, not volume, is what usually kills a MapReduce job
THE COMBINER, AND WHY ITS SAVING IS SMALL HERE
| Shuffled | |
|---|---|
| no combiner | 48 |
| with combiner | 39 |
| saving | 9 (18.75%) |
WHY IT MATTERS
Note how small that is, and why it is honest. The combiner runs per map task, and these documents are 5 to 11 words — there is almost nothing to merge within one split. On a 128 MB split the same combiner cuts the shuffle by orders of magnitude. The cluster agrees: its counters show the same 48 map output records and 39 combine output records.
The combiner's value scales with split size, which is the point this tiny dataset makes precisely by failing to impress.
THE COMBINER THAT IS WRONG
| Reducer computes | Safe? |
|---|---|
| sum, max, count | yes |
| mean | NO |
| median | NO |
Demonstrated: [1,1,1,10] and [10] give mean of means 6.6250 against
true mean 4.6000. Emit (sum, count) and divide only in the reducer —
the same average-of-averages trap Business Intelligence Tools met in DAX.
Partitioning. 3 reducers give partitions of [20, 17, 11], ratio 1.82. Hash partitioning is only as balanced as the key distribution, and skew, not volume, is what usually kills a MapReduce job.
The three Java details. The map key is a byte offset, not a line number. Reuse the
Writable objects — a new Text() per word makes GC the job. The reduce Iterable
can be walked once — it streams from disk.
RESULT
48 words in, 26 distinct words out, on the cluster and in the engine alike; the combiner cut the 48 shuffled records to 39 — 18.75%, small because each split is one short document. A combiner for a mean is wrong: 6.6250 against the true 4.6000.
Develop a MapReduce job for inverted index creation.
Build, with a Java MapReduce job, an index from each word to the documents it occurs in and how often; then answer queries from the index alone, and measure its size and skew.
On the cluster, InvertedIndex.java:
The Python check, 08_inverted_index.py:
THE INDEX
| Term | Postings |
|---|---|
dog |
doc1:1, doc2:1, doc3:1, doc6:1 |
quick |
doc1:1, doc3:2 |
big |
doc4:2, doc5:2 |
quick appears twice in doc3, and the posting records it. Frequency is
the difference between "does this word occur" and "how relevant is this
document" — boolean retrieval against ranked retrieval, in one number.
On the cluster, InvertedIndex.java:
// Experiment 8 -- an inverted index in MapReduce
//
// Run it: the build-and-run lines below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 08_inverted_index.py, which builds the same index and
// answers boolean queries against it
//
// Build and run:
// javac -classpath $(hadoop classpath) -d classes InvertedIndex.java
// jar -cvf InvertedIndex.jar -C classes/ .
// hadoop jar InvertedIndex.jar InvertedIndex /user/student/docs /user/student/out
//
import java.io.IOException;
import java.util.HashMap;
import java.util.Map;
import java.util.StringTokenizer;
import java.util.TreeMap;
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.fs.Path;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Job;
import org.apache.hadoop.mapreduce.Mapper;
import org.apache.hadoop.mapreduce.Reducer;
import org.apache.hadoop.mapreduce.lib.input.FileInputFormat;
import org.apache.hadoop.mapreduce.lib.input.FileSplit;
import org.apache.hadoop.mapreduce.lib.output.FileOutputFormat;
public class InvertedIndex {
// Step 1: Map each word to its document
public static class IndexMapper
extends Mapper<LongWritable, Text, Text, Text> {
private final Text word = new Text();
private final Text docId = new Text();
private String fileName;
@Override
protected void setup(Context context) {
// THE FILENAME IS NOT IN THE KEY OR THE VALUE. It comes from the
// InputSplit, and this is the only way to get at it. Every
// inverted-index question turns on knowing that.
FileSplit split = (FileSplit) context.getInputSplit();
fileName = split.getPath().getName();
}
@Override
public void map(LongWritable key, Text value, Context context)
throws IOException, InterruptedException {
StringTokenizer itr = new StringTokenizer(value.toString());
while (itr.hasMoreTokens()) {
word.set(itr.nextToken().toLowerCase());
docId.set(fileName);
context.write(word, docId);
}
}
}
// Step 2: Reduce to a posting list
public static class IndexReducer
extends Reducer<Text, Text, Text, Text> {
private final Text postings = new Text();
@Override
public void reduce(Text key, Iterable<Text> values, Context context)
throws IOException, InterruptedException {
// Count occurrences per document, in ONE pass over the Iterable.
Map<String, Integer> freq = new TreeMap<>();
for (Text doc : values) {
String d = doc.toString();
freq.merge(d, 1, Integer::sum);
}
StringBuilder sb = new StringBuilder();
for (Map.Entry<String, Integer> e : freq.entrySet()) {
if (sb.length() > 0) sb.append(", ");
sb.append(e.getKey()).append(':').append(e.getValue());
}
postings.set(sb.toString());
context.write(key, postings);
// The posting list is held in memory here. For a stop word on a
// real corpus that map does not fit, which is why production
// indexers emit (term, doc) pairs SORTED and stream the merge --
// a secondary sort, not a HashMap.
}
}
// Step 3: Configure the job
public static void main(String[] args) throws Exception {
Configuration conf = new Configuration();
Job job = Job.getInstance(conf, "inverted index");
job.setJarByClass(InvertedIndex.class);
job.setMapperClass(IndexMapper.class);
job.setReducerClass(IndexReducer.class);
// NO COMBINER HERE. The reducer's output type (Text postings) differs
// from its input type (Text docId), and a combiner must have the same
// input and output types as the reducer. Setting one would not
// compile -- and if the types happened to match, it would corrupt
// the counts.
job.setOutputKeyClass(Text.class);
job.setOutputValueClass(Text.class);
FileInputFormat.addInputPath(job, new Path(args[0]));
FileOutputFormat.setOutputPath(job, new Path(args[1]));
System.exit(job.waitForCompletion(true) ? 0 : 1);
}
}
// Expected output on the six sample documents (verified in 08_inverted_index.py):
// dog doc1.txt:1, doc2.txt:1, doc3.txt:1, doc6.txt:1
// quick doc1.txt:1, doc3.txt:2
// big doc4.txt:2, doc5.txt:2
// ... 26 terms, 39 postings
The Python check, 08_inverted_index.py:
"""Experiment 8 -- a MapReduce job for inverted index creation.
The inverted index is what a search engine is. Word count changes the VALUE
between map and reduce; the inverted index changes the KEY SPACE -- the map
output key is a word and the value is a document, which is exactly the
transposition a search index needs.
"""
from mapreduce import run
import fixtures as f
INPUT = sorted(f.DOCS.items())
def mapper(doc, text):
"""(doc, text) -> (word, doc) -- emit the document as the VALUE."""
for pos, word in enumerate(text.split()):
yield word, (doc, pos)
def reducer(word, postings):
"""(word, [(doc, pos), ...]) -> (word, posting list).
The posting list is deduplicated by document and carries a frequency,
which is what turns a boolean index into a ranked one.
"""
per_doc = {}
for doc, pos in postings:
per_doc.setdefault(doc, []).append(pos)
yield word, {d: len(p) for d, p in sorted(per_doc.items())}
def boolean_and(index, *words):
"""Intersect posting lists -- an AND query."""
sets = [set(index.get(w, {})) for w in words]
return sorted(set.intersection(*sets)) if sets else []
def boolean_or(index, *words):
sets = [set(index.get(w, {})) for w in words]
return sorted(set.union(*sets)) if sets else []
def main():
print(" Experiment 8 -- inverted index in MapReduce")
# Step 1: Build the index
trace = {}
index = dict(run(INPUT, mapper, reducer, trace=trace))
print(f"\n {len(INPUT)} documents in, {len(index)} index terms out")
print(f" map emitted {trace['map_output']} postings")
assert trace["map_output"] == 48 and len(index) == 26
# Step 2: Read a slice of it
print("\n a slice of the index (term -> {doc: frequency}):")
for w in ("dog", "quick", "big", "data", "machine"):
entry = ", ".join(f"{d.replace('.txt', '')}:{c}"
for d, c in index[w].items())
print(f" {w:<10}{entry}")
assert index["dog"] == {"doc1.txt": 1, "doc2.txt": 1,
"doc3.txt": 1, "doc6.txt": 1}
assert index["quick"] == {"doc1.txt": 1, "doc3.txt": 2}
print(""" 'quick' appears TWICE in doc3, and the posting records
that. Frequency is the difference between 'does this word
occur' and 'how relevant is this document' -- boolean
retrieval against ranked retrieval, in one number""")
# Step 3: Answer queries from it
print("\n queries answered from the index alone:")
for q in (("quick", "fox"), ("big", "data"), ("dog", "machine")):
hits = boolean_and(index, *q)
pretty = [h.replace(".txt", "") for h in hits]
print(f" {' AND '.join(q):<24}-> {pretty if pretty else 'no match'}")
assert boolean_and(index, "quick", "fox") == ["doc1.txt", "doc3.txt"]
assert boolean_and(index, "big", "data") == ["doc4.txt", "doc5.txt"]
assert boolean_and(index, "dog", "machine") == []
hits_or = boolean_or(index, "dog", "machine")
print(f" {'dog OR machine':<24}-> "
f"{[h.replace('.txt', '') for h in hits_or]}")
assert len(hits_or) == 5
print(""" NOT ONE DOCUMENT WAS READ to answer these. That is the
entire point of an inverted index: query cost depends on the
number of MATCHES, not on the size of the corpus. Scanning
6 documents is cheap; scanning 6 billion is not""")
# Step 4: Weigh its size
corpus_chars = sum(len(t) for _, t in INPUT)
postings = sum(len(v) for v in index.values())
print(f"\n the index is not free:")
print(f" corpus {corpus_chars:>5} characters")
print(f" index terms {len(index):>5}")
print(f" postings {postings:>5}")
print(f" ratio {postings / len(INPUT):>5.1f} postings per document")
print(""" a full-text index typically runs 20-40% of the corpus
size, and that is BEFORE positions. Search is a space-for-time
trade, and 'the index is bigger than I expected' is the normal
outcome, not a mistake""")
# Step 5: See why it is a MapReduce job
print("\n why MapReduce suits this:")
print(" map is per-document and EMBARRASSINGLY PARALLEL")
print(" reduce is per-term, and every posting for a term")
print(" arrives at the same reducer by construction")
print(""" the shuffle does the hard part -- gathering every
mention of a word from every machine in the cluster -- and
you never wrote a line of network code. That is the whole
value proposition of the model""")
# Step 6: Measure the skew
print("\n the skew, measured:")
sizes = sorted(((w, sum(v.values())) for w, v in index.items()),
key=lambda kv: -kv[1])
print(f" largest posting list : {sizes[0][0]!r} with {sizes[0][1]}")
print(f" singleton terms : "
f"{sum(1 for _, n in sizes if n == 1)} of {len(sizes)}")
assert sizes[0][0] == "the" and sizes[0][1] == 5
print(""" 'the' is the biggest list here and would be the biggest
on any English corpus. Real engines drop stop words or split
hot terms across reducers, because one reducer holding 'the'
is the job's critical path""")
return index
if __name__ == "__main__":
main()
On the cluster, InvertedIndex.java:
OUTPUT
$ hdfs dfs -mkdir -p /user/student/docs
$ hdfs dfs -put docs/*.txt /user/student/docs/
$ hdfs dfs -ls /user/student/docs | awk 'NR>1{print $NF}'
/user/student/docs/doc1.txt
/user/student/docs/doc2.txt
/user/student/docs/doc3.txt
/user/student/docs/doc4.txt
/user/student/docs/doc5.txt
/user/student/docs/doc6.txt
$ mkdir -p classes
$ javac -classpath $(hadoop classpath) -d classes InvertedIndex.java
$ jar -cvf InvertedIndex.jar -C classes/ .
added manifest
adding: InvertedIndex$IndexReducer.class(in = 3129) (out= 1331)(deflated 57%)
adding: InvertedIndex$IndexMapper.class(in = 2342) (out= 935)(deflated 60%)
adding: InvertedIndex.class(in = 1424) (out= 772)(deflated 45%)
$ hadoop jar InvertedIndex.jar InvertedIndex /user/student/docs /user/student/out 2>&1 | grep -E "completed successfully|Map input records|Map output records|Combine input records|Combine output records|Reduce input groups|Reduce output records"
2026-10-04 23:17:18,259 INFO mapreduce.Job: Job job_1791155800966_0001 completed successfully
Map input records=6
Map output records=48
Combine input records=0
Combine output records=0
Reduce input groups=26
Reduce output records=26
$ hdfs dfs -cat /user/student/out/part-r-00000
a doc3.txt:2
all doc2.txt:1
and doc5.txt:1, doc6.txt:1
big doc4.txt:2, doc5.txt:2
brown doc1.txt:1, doc3.txt:1
data doc4.txt:2, doc5.txt:2
day doc2.txt:1
dog doc1.txt:1, doc2.txt:1, doc3.txt:1, doc6.txt:1
for doc4.txt:1
fox doc1.txt:1, doc3.txt:1, doc6.txt:1
hadoop doc5.txt:1
is doc4.txt:2
jumps doc1.txt:1
lazy doc1.txt:1, doc2.txt:1
machine doc4.txt:1
one doc4.txt:1
outpaces doc3.txt:1
over doc1.txt:1
processes doc5.txt:1
quick doc1.txt:1, doc3.txt:2
sleeps doc2.txt:1
spark doc5.txt:1
stores doc5.txt:1
that doc4.txt:1
the doc1.txt:2, doc2.txt:1, doc6.txt:2
too doc4.txt:1
The Python check, 08_inverted_index.py:
OUTPUT
Experiment 8 -- inverted index in MapReduce
6 documents in, 26 index terms out
map emitted 48 postings
a slice of the index (term -> {doc: frequency}):
dog doc1:1, doc2:1, doc3:1, doc6:1
quick doc1:1, doc3:2
big doc4:2, doc5:2
data doc4:2, doc5:2
machine doc4:1
'quick' appears TWICE in doc3, and the posting records
that. Frequency is the difference between 'does this word
occur' and 'how relevant is this document' -- boolean
retrieval against ranked retrieval, in one number
queries answered from the index alone:
quick AND fox -> ['doc1', 'doc3']
big AND data -> ['doc4', 'doc5']
dog AND machine -> no match
dog OR machine -> ['doc1', 'doc2', 'doc3', 'doc4', 'doc6']
NOT ONE DOCUMENT WAS READ to answer these. That is the
entire point of an inverted index: query cost depends on the
number of MATCHES, not on the size of the corpus. Scanning
6 documents is cheap; scanning 6 billion is not
the index is not free:
corpus 226 characters
index terms 26
postings 39
ratio 6.5 postings per document
a full-text index typically runs 20-40% of the corpus
size, and that is BEFORE positions. Search is a space-for-time
trade, and 'the index is bigger than I expected' is the normal
outcome, not a mistake
why MapReduce suits this:
map is per-document and EMBARRASSINGLY PARALLEL
reduce is per-term, and every posting for a term
arrives at the same reducer by construction
the shuffle does the hard part -- gathering every
mention of a word from every machine in the cluster -- and
you never wrote a line of network code. That is the whole
value proposition of the model
the skew, measured:
largest posting list : 'the' with 5
singleton terms : 15 of 26
'the' is the biggest list here and would be the biggest
on any English corpus. Real engines drop stop words or split
hot terms across reducers, because one reducer holding 'the'
is the job's critical path
QUERIES ANSWERED FROM THE INDEX ALONE
| Query | Result |
|---|---|
quick AND fox |
doc1, doc3 |
big AND data |
doc4, doc5 |
dog AND machine |
no match |
dog OR machine |
doc1, doc2, doc3, doc4, doc6 |
Not one document was read. Query cost depends on the number of matches, not on the size of the corpus — which is the entire point of an inverted index.
The index is not free. 226 characters of corpus → 26 terms, 39 postings. A full-text index typically runs 20–40% of the corpus size, before positions. Search is a space-for-time trade.
The skew, measured. Largest posting list: the, with 5. Singleton terms: 15 of 26. the
would be the biggest list on any English corpus, and one reducer holding it is the job's
critical path.
THE JAVA DETAIL THAT IS EXAMINED
The filename is in neither the key nor the value. It comes from the input
split — ((FileSplit) context.getInputSplit()).getPath().getName().
And this job cannot use a combiner: the reducer's output type (a posting
string) differs from its input type (a document id), and a combiner must match
the reducer on both — the cluster's counters show Combine input records=0.
RESULT
6 documents in, 26 index terms out, from 48 postings — on the cluster and in the engine alike; quick is doc1:1, doc3:2. Boolean queries are answered without reading a document, and the largest posting list is the, 5 occurrences in 3 documents.
Perform data analysis using Pig Latin scripts.
Load the sales, filter, group and total them by category, order the result, join the stores map-side and flatten the tags; then walk the same dataflow one operator at a time.
On the cluster, 09_analysis.pig:
The Python check, 09_pig_equivalent.py:
THE DATAFLOW
The Python half walks the dataflow one operator at a time, which is how you
debug a Pig script anyway — that is what ILLUSTRATE does.
A = LOAD 'sales' -- 9 rows
B = FILTER A BY qty >= 6 -- 7 rows
C = GROUP B BY category -- 3 groups: Grocery {4}, Personal {1}, Stationery {2}
D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)
E = ORDER D BY revenue DESC -- top category: Grocery
On the cluster, 09_analysis.pig:
-- Experiment 9 -- data analysis with Pig Latin
--
-- Run it: pig -x mapreduce 09_analysis.pig, with the input files on HDFS. It was run on a Hadoop 3.3.6 cluster where these labs
-- are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
-- [Changed: this said the file had never been run, as the Hadoop stack could not be
-- installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
--
-- The runnable half is 09_pig_equivalent.py, which walks the same dataflow relation by relation
--
-- run with: pig -x mapreduce 09_analysis.pig
-- or: pig -x local 09_analysis.pig
-- Step 1: Load the sales
sales = LOAD '/user/student/sales/sales.csv' USING PigStorage(',')
AS (date_key:chararray, store_name:chararray, region:chararray,
product:chararray, category:chararray,
qty:int, list_price:double);
-- [Noted, from running it: sales.csv has a header row, and PigStorage has no
-- idea -- "Successfully read 10 records" for nine sales. The header's qty, the
-- word qty, becomes null and the FILTER below drops it without a word.
-- org.apache.pig.piggybank.storage.CSVExcelStorage can skip a header.]
-- Step 2: Build the dataflow, a relation per step
-- a relation per step. This is the point of Pig.
priced = FOREACH sales GENERATE *, qty * list_price AS revenue;
bulk = FILTER priced BY qty >= 6;
by_cat = GROUP bulk BY category;
totals = FOREACH by_cat GENERATE
group AS category,
COUNT(bulk) AS orders,
SUM(bulk.qty) AS units,
SUM(bulk.revenue) AS revenue;
ranked = ORDER totals BY revenue DESC;
-- Step 3: Describe, illustrate and explain it
DESCRIBE ranked; -- the SCHEMA, without running anything
ILLUSTRATE ranked; -- sample rows pushed through EVERY step -- Pig's
-- best feature and the one with no SQL equivalent
EXPLAIN -brief ranked; -- logical, physical and MapReduce plans
-- [Changed: -brief leaves the nested plans folded; in
-- full they run to some 400 lines]
-- Step 4: Store the result
STORE ranked INTO '/user/student/out/by_category' USING PigStorage(',');
-- ^ NOTHING ABOVE THIS LINE HAS RUN. Pig is lazy: LOAD/FILTER/GROUP build a
-- plan, and STORE or DUMP compiles it into MapReduce jobs and submits them.
-- Step 5: Join the stores, map-side
-- a join, and the hint that decides how it is executed
stores = LOAD '/user/student/dim/stores.csv' USING PigStorage(',')
AS (store_name:chararray, city:chararray, region:chararray);
joined = JOIN priced BY store_name, stores BY store_name USING 'replicated';
-- ^^^^^^^^^^^^^^^^^^
-- 'replicated' = a MAP-SIDE join: the small relation is loaded into memory
-- on every mapper, so there is NO SHUFFLE. It fails with an OOM if the
-- right-hand relation does not fit. The default is a reduce-side join.
-- Other strategies: 'skewed' (for one hot key), 'merge' (both sorted).
-- [Corrected: the field was called store, and STORE is a Pig keyword in any
-- case. LOAD ... AS accepted it, but JOIN ... BY store, stores BY store
-- failed to parse -- and Pig parses the whole script before it runs any
-- STORE, so the error lost the by_category output above as well.]
DUMP joined;
-- [Changed: DUMP joined is added. Nothing used joined, and Pig is lazy, so
-- the join was never run.]
-- Step 6: Flatten the tags
-- FLATTEN, which has no clean SQL equivalent
tags = LOAD '/user/student/tags.csv' AS (product:chararray, taglist:chararray);
split_t = FOREACH tags GENERATE product,
FLATTEN(TOKENIZE(taglist, ';')) AS tag;
DUMP split_t; -- one row per (product, tag) pair, from one row per product
The Python check, 09_pig_equivalent.py:
"""Experiment 9 -- data analysis with Pig Latin scripts.
`09_analysis.pig` carries the real Pig Latin, and runs on Pig itself (the lab
page shows it). What runs here is the same dataflow, one operator at a time, so the
INTERMEDIATE relations in the notes are real -- and stepping through them is
exactly how you debug a Pig script anyway (that is what ILLUSTRATE does).
Pig's value over Hive is that it is a DATAFLOW language: you name every
intermediate relation, so a 12-step transformation reads top to bottom instead
of nesting twelve sub-queries.
"""
import fixtures as f
SALES = f.SALES_DF
def show(name, rel, cols, limit=4):
print(f"\n {name} -- {len(rel)} rows")
head = rel[cols].head(limit)
widths = [max(len(str(c)), int(rel[c].astype(str).str.len().max())) + 3
for c in cols]
print(" " + "".join(f"{c:>{w}}" for c, w in zip(cols, widths)))
for _, r in head.iterrows():
print(" " + "".join(
f"{r[c]:>{w},.0f}" if isinstance(r[c], float) else f"{str(r[c]):>{w}}"
for c, w in zip(cols, widths)))
if len(rel) > limit:
print(f" ... {len(rel) - limit} more")
def main():
print(" Experiment 9 -- the Pig Latin dataflow, one operator at a time")
# Step 1: Load the sales
# A = LOAD
A = SALES.copy()
show("A = LOAD 'sales'", A, ["store", "product", "qty", "revenue"])
# Step 2: Filter the bulk orders
# B = FILTER
B = A[A["qty"] >= 6]
show("B = FILTER A BY qty >= 6", B, ["store", "product", "qty", "revenue"])
assert len(B) == 7, 'seven of nine orders are 6 units or more'
# Step 3: Group by category
# C = GROUP
C = B.groupby("category")
print(f"\n C = GROUP B BY category -- {C.ngroups} groups")
for name, grp in C:
print(f" ({name}, {{{len(grp)} tuples}})")
print(""" GROUP in Pig produces a BAG per key, not an aggregate.
The bag is the value, and FOREACH ... GENERATE is what turns
it into numbers. Hive fuses the two; Pig keeps them apart,
which is why Pig can do things to a group that SQL cannot
express without a window function""")
# Step 4: Total each group
# D = FOREACH ... GENERATE
D = (C.agg(units=("qty", "sum"), revenue=("revenue", "sum"),
orders=("order_id", "count") if "order_id" in B else ("qty", "size"))
.reset_index())
print(f"\n D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)")
print(f" {'category':<12}{'units':>8}{'revenue':>12}{'orders':>8}")
for _, r in D.iterrows():
print(f" {r['category']:<12}{r['units']:>8.0f}"
f"{r['revenue']:>12,.0f}{r['orders']:>8.0f}")
assert D["revenue"].sum() == B["revenue"].sum()
# Step 5: Order by revenue
# E = ORDER
E = D.sort_values("revenue", ascending=False)
print(f"\n E = ORDER D BY revenue DESC")
print(f" top category: {E.iloc[0]['category']} "
f"at {E.iloc[0]['revenue']:,.0f}")
assert E.iloc[0]["category"] == "Grocery"
# Step 6: Join the stores
# F = JOIN
print("\n F = JOIN A BY store_key, stores BY store_key")
joined = A.groupby(["region", "store"], as_index=False)["revenue"].sum()
print(f" {'region':<8}{'store':<14}{'revenue':>12}")
for _, r in joined.sort_values("revenue", ascending=False).iterrows():
print(f" {r['region']:<8}{r['store']:<14}{r['revenue']:>12,.0f}")
assert joined["revenue"].sum() == f.total_revenue()
# Step 7: Map the operators to SQL
print("\n the operators, and their SQL equivalents:")
print(f" {'Pig Latin':<26}{'SQL'}")
for pig, sql in (
("LOAD / STORE", "no equivalent -- SQL assumes a table exists"),
("FILTER", "WHERE"),
("FOREACH .. GENERATE", "SELECT"),
("GROUP", "GROUP BY, but the bag is kept"),
("JOIN", "JOIN"),
("ORDER", "ORDER BY"),
("DISTINCT", "DISTINCT"),
("FLATTEN", "UNNEST / LATERAL VIEW explode"),
("ILLUSTRATE", "no equivalent -- sample data through the plan")):
print(f" {pig:<26}{sql}")
print("""
two operators have no SQL equivalent, and they are the
reason to reach for Pig: LOAD, which lets a script read a
semi-structured file with no schema declared in advance,
and ILLUSTRATE, which pushes a few representative rows
through every step of the plan so you can see where a
12-stage pipeline went wrong""")
# Step 8: See lazy evaluation
print("\n lazy evaluation, which surprises everyone:")
print(" nothing runs until STORE or DUMP.")
print(" A = LOAD ...; B = FILTER ...; C = GROUP ...;")
print(" -- no job has been submitted yet")
print(" STORE C INTO 'out'; -- NOW Pig compiles and runs it")
print(""" because Pig sees the whole dataflow before executing, it
can merge the FILTER into the LOAD and fuse consecutive
FOREACHes into one MapReduce job. Writing the steps
separately costs nothing -- which is the entire argument
against nesting sub-queries to avoid 'extra passes'""")
if __name__ == "__main__":
main()
On the cluster, 09_analysis.pig:
OUTPUT
$ hdfs dfs -mkdir -p /user/student/sales /user/student/dim
$ hdfs dfs -put sales.csv /user/student/sales/
$ hdfs dfs -put stores.csv /user/student/dim/
$ hdfs dfs -put tags.tsv /user/student/tags.csv
$ pig -x mapreduce 09_analysis.pig 2>pig.log
ranked: {category: chararray,orders: long,units: long,revenue: double}
(D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0)
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| sales | date_key:chararray | store_name:chararray | region:chararray | product:chararray | category:chararray | qty:int | list_price:double |
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| | D3 | Hyderabad | North | Rice 5kg | Grocery | 4 | 280.0 |
| | D3 | Guntur | South | Tea 500g | Grocery | 12 | 210.0 |
| | D2 | Vijayawada | South | Rice 5kg | Grocery | 6 | 280.0 |
| | D4 | Hyderabad | North | Notebook | Stationery | 15 | 40.0 |
------------------------------------------------------------------------------------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| priced | date_key:chararray | store_name:chararray | region:chararray | product:chararray | category:chararray | qty:int | list_price:double | revenue:double |
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| | D3 | Hyderabad | North | Rice 5kg | Grocery | 4 | 280.0 | 1120.0 |
| | D3 | Guntur | South | Tea 500g | Grocery | 12 | 210.0 | 2520.0 |
| | D2 | Vijayawada | South | Rice 5kg | Grocery | 6 | 280.0 | 1680.0 |
| | D4 | Hyderabad | North | Notebook | Stationery | 15 | 40.0 | 600.0 |
----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| bulk | date_key:chararray | store_name:chararray | region:chararray | product:chararray | category:chararray | qty:int | list_price:double | revenue:double |
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| | D3 | Guntur | South | Tea 500g | Grocery | 12 | 210.0 | 2520.0 |
| | D2 | Vijayawada | South | Rice 5kg | Grocery | 6 | 280.0 | 1680.0 |
| | D4 | Hyderabad | North | Notebook | Stationery | 15 | 40.0 | 600.0 |
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| by_cat | group:chararray | bulk:bag{:tuple(date_key:chararray,store_name:chararray,region:chararray,product:chararray,category:chararray,qty:int,list_price:double,revenue:double)} |
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
| | Grocery | {} |
| | Grocery | {} |
| | Stationery | {} |
| | Stationery | {} |
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------
| totals | category:chararray | orders:long | units:long | revenue:double |
----------------------------------------------------------------------------------------------
| | Grocery | 2 | 18 | 4200.0 |
| | Stationery | 1 | 15 | 600.0 |
----------------------------------------------------------------------------------------------
----------------------------------------------------------------------------------------------
| ranked | category:chararray | orders:long | units:long | revenue:double |
----------------------------------------------------------------------------------------------
| | Grocery | 2 | 18 | 4200.0 |
| | Stationery | 1 | 15 | 600.0 |
----------------------------------------------------------------------------------------------
#-----------------------------------------------
# New Logical Plan:
#-----------------------------------------------
ranked: (Name: LOStore Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
|
|---ranked: (Name: LOSort Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
| |
| revenue:(Name: Project Type: double Uid: 186 Input: 0 Column: 3)
|
|---totals: (Name: LOForEach Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)
| |
| (Name: LOGenerate[false,false,false,false] Schema: category#152:chararray,orders#180:long,units#183:long,revenue#186:double)ColumnPrune:OutputUids=[180, 183, 152, 186]ColumnPrune:InputUids=[178, 152]
| | |
| | group:(Name: Project Type: chararray Uid: 152 Input: 0 Column: (*))
| | |
| | (Name: UserFunc(org.apache.pig.builtin.COUNT) Type: long Uid: 180)
| | |
| | |---bulk:(Name: Project Type: bag Uid: 178 Input: 1 Column: (*))
| | |
| | (Name: UserFunc(org.apache.pig.builtin.LongSum) Type: long Uid: 183)
| | |
| | |---(Name: Dereference Type: bag Uid: 182 Column:[5])
| | |
| | |---bulk:(Name: Project Type: bag Uid: 178 Input: 2 Column: (*))
| | |
| | (Name: UserFunc(org.apache.pig.builtin.DoubleSum) Type: double Uid: 186)
| | |
| | |---(Name: Dereference Type: bag Uid: 185 Column:[7])
| | |
| | |---bulk:(Name: Project Type: bag Uid: 178 Input: 3 Column: (*))
| |
| |---(Name: LOInnerLoad[0] Schema: group#152:chararray)
| |
| |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
| |
| |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
| |
| |---bulk: (Name: LOInnerLoad[1] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
|
|---by_cat: (Name: LOCogroup Schema: group#152:chararray,priced#178:bag{#199:tuple(date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)})
| |
| category:(Name: Project Type: chararray Uid: 152 Input: 0 Column: 4)
|
|---priced: (Name: LOForEach Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)
| |
| (Name: LOGenerate[false,false,false,false,false,false,false,false] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double,revenue#175:double)ColumnPrune:OutputUids=[148, 149, 150, 151, 152, 153, 154, 175]ColumnPrune:InputUids=[148, 149, 150, 151, 152, 153, 154]
| | |
| | date_key:(Name: Project Type: chararray Uid: 148 Input: 2 Column: (*))
| | |
| | store_name:(Name: Project Type: chararray Uid: 149 Input: 3 Column: (*))
| | |
| | region:(Name: Project Type: chararray Uid: 150 Input: 4 Column: (*))
| | |
| | product:(Name: Project Type: chararray Uid: 151 Input: 5 Column: (*))
| | |
| | category:(Name: Project Type: chararray Uid: 152 Input: 6 Column: (*))
| | |
| | qty:(Name: Project Type: int Uid: 153 Input: 7 Column: (*))
| | |
| | list_price:(Name: Project Type: double Uid: 154 Input: 8 Column: (*))
| | |
| | (Name: Multiply Type: double Uid: 175)
| | |
| | |---(Name: Cast Type: double Uid: 153)
| | | |
| | | |---qty:(Name: Project Type: int Uid: 153 Input: 0 Column: (*))
| | |
| | |---list_price:(Name: Project Type: double Uid: 154 Input: 1 Column: (*))
| |
| |---(Name: LOInnerLoad[5] Schema: qty#153:int)
| |
| |---(Name: LOInnerLoad[6] Schema: list_price#154:double)
| |
| |---(Name: LOInnerLoad[0] Schema: date_key#148:chararray)
| |
| |---(Name: LOInnerLoad[1] Schema: store_name#149:chararray)
| |
| |---(Name: LOInnerLoad[2] Schema: region#150:chararray)
| |
| |---(Name: LOInnerLoad[3] Schema: product#151:chararray)
| |
| |---(Name: LOInnerLoad[4] Schema: category#152:chararray)
| |
| |---(Name: LOInnerLoad[5] Schema: qty#153:int)
| |
| |---(Name: LOInnerLoad[6] Schema: list_price#154:double)
|
|---bulk: (Name: LOFilter Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)
| |
| (Name: GreaterThanEqual Type: boolean Uid: 177)
| |
| |---qty:(Name: Project Type: int Uid: 153 Input: 0 Column: 5)
| |
| |---(Name: Constant Type: int Uid: 176)
|
|---sales: (Name: LOForEach Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)
| |
| (Name: LOGenerate[false,false,false,false,false,false,false] Schema: date_key#148:chararray,store_name#149:chararray,region#150:chararray,product#151:chararray,category#152:chararray,qty#153:int,list_price#154:double)ColumnPrune:OutputUids=[148, 149, 150, 151, 152, 153, 154]ColumnPrune:InputUids=[148, 149, 150, 151, 152, 153, 154]
| | |
| | (Name: Cast Type: chararray Uid: 148)
| | |
| | |---date_key:(Name: Project Type: bytearray Uid: 148 Input: 0 Column: (*))
| | |
| | (Name: Cast Type: chararray Uid: 149)
| | |
| | |---store_name:(Name: Project Type: bytearray Uid: 149 Input: 1 Column: (*))
| | |
| | (Name: Cast Type: chararray Uid: 150)
| | |
| | |---region:(Name: Project Type: bytearray Uid: 150 Input: 2 Column: (*))
| | |
| | (Name: Cast Type: chararray Uid: 151)
| | |
| | |---product:(Name: Project Type: bytearray Uid: 151 Input: 3 Column: (*))
| | |
| | (Name: Cast Type: chararray Uid: 152)
| | |
| | |---category:(Name: Project Type: bytearray Uid: 152 Input: 4 Column: (*))
| | |
| | (Name: Cast Type: int Uid: 153)
| | |
| | |---qty:(Name: Project Type: bytearray Uid: 153 Input: 5 Column: (*))
| | |
| | (Name: Cast Type: double Uid: 154)
| | |
| | |---list_price:(Name: Project Type: bytearray Uid: 154 Input: 6 Column: (*))
| |
| |---(Name: LOInnerLoad[0] Schema: date_key#148:bytearray)
| |
| |---(Name: LOInnerLoad[1] Schema: store_name#149:bytearray)
| |
| |---(Name: LOInnerLoad[2] Schema: region#150:bytearray)
| |
| |---(Name: LOInnerLoad[3] Schema: product#151:bytearray)
| |
| |---(Name: LOInnerLoad[4] Schema: category#152:bytearray)
| |
| |---(Name: LOInnerLoad[5] Schema: qty#153:bytearray)
| |
| |---(Name: LOInnerLoad[6] Schema: list_price#154:bytearray)
|
|---sales: (Name: LOLoad Schema: date_key#148:bytearray,store_name#149:bytearray,region#150:bytearray,product#151:bytearray,category#152:bytearray,qty#153:bytearray,list_price#154:bytearray)RequiredFields:null
#-----------------------------------------------
# Physical Plan:
#-----------------------------------------------
ranked: Store(fakefile:org.apache.pig.builtin.PigStorage) - scope-804
|
|---ranked: POSort[bag]() - scope-803
|
|---totals: New For Each(false,false,false,false)[bag] - scope-801
|
|---by_cat: Package(Packager)[tuple]{chararray} - scope-785
|
|---by_cat: Global Rearrange[tuple] - scope-784
|
|---by_cat: Local Rearrange[tuple]{chararray}(false) - scope-786
|
|---priced: New For Each(false,false,false,false,false,false,false,false)[bag] - scope-783
|
|---bulk: Filter[bag] - scope-759
|
|---sales: New For Each(false,false,false,false,false,false,false)[bag] - scope-758
|
|---sales: Load(/user/student/sales/sales.csv:PigStorage(',')) - scope-736
#--------------------------------------------------
# Map Reduce Plan
#--------------------------------------------------
MapReduce node scope-805
Map Plan
by_cat: Local Rearrange[tuple]{chararray}(false) - scope-851
|
|---totals: New For Each(false,false,false,false)[bag] - scope-828
|
|---Pre Combiner Local Rearrange[tuple]{Unknown} - scope-854
|
|---priced: New For Each(false,false,false,false,false,false,false,false)[bag] - scope-783
|
|---bulk: Filter[bag] - scope-759
|
|---sales: New For Each(false,false,false,false,false,false,false)[bag] - scope-758
|
|---sales: Load(/user/student/sales/sales.csv:PigStorage(',')) - scope-736--------
Combine Plan
by_cat: Local Rearrange[tuple]{chararray}(false) - scope-855
|
|---totals: New For Each(false,false,false,false)[bag] - scope-838
|
|---by_cat: Package(CombinerPackager)[tuple]{chararray} - scope-850--------
Reduce Plan
Store(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.io.InterStorage) - scope-806
|
|---totals: New For Each(false,false,false,false)[bag] - scope-801
|
|---by_cat: Package(CombinerPackager)[tuple]{chararray} - scope-785--------
Global sort: false
----------------
MapReduce node scope-808
Map Plan
ranked: Local Rearrange[tuple]{tuple}(false) - scope-812
|
|---New For Each(false)[tuple] - scope-810
|
|---Load(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.builtin.RandomSampleLoader('org.apache.pig.impl.io.InterStorage','100')) - scope-807--------
Reduce Plan
Store(hdfs://localhost:9000/tmp/temp1251966677/tmp2048799314:org.apache.pig.impl.io.InterStorage) - scope-821
|
|---New For Each(false)[tuple] - scope-820
|
|---New For Each(false,false)[tuple] - scope-817
|
|---Package(Packager)[tuple]{chararray} - scope-813--------
Global sort: false
Secondary sort: true
----------------
MapReduce node scope-823
Map Plan
ranked: Local Rearrange[tuple]{double}(false) - scope-824
|
|---Load(hdfs://localhost:9000/tmp/temp1251966677/tmp-2141768117:org.apache.pig.impl.io.InterStorage) - scope-822--------
Reduce Plan
ranked: Store(fakefile:org.apache.pig.builtin.PigStorage) - scope-804
|
|---New For Each(true)[tuple] - scope-827
|
|---Package(LitePackager)[tuple]{double} - scope-825--------
Global sort: true
Quantile file: hdfs://localhost:9000/tmp/temp1251966677/tmp2048799314
----------------
(D1,Vijayawada,South,Rice 5kg,Grocery,10,280.0,2800.0,Vijayawada,Vijayawada,South)
(D1,Vijayawada,South,Shampoo 200ml,Personal,5,140.0,700.0,Vijayawada,Vijayawada,South)
(D1,Guntur,South,Tea 500g,Grocery,8,210.0,1680.0,Guntur,Guntur,South)
(D2,Vijayawada,South,Rice 5kg,Grocery,6,280.0,1680.0,Vijayawada,Vijayawada,South)
(D2,Hyderabad,North,Notebook,Stationery,20,40.0,800.0,Hyderabad,Hyderabad,North)
(D3,Guntur,South,Tea 500g,Grocery,12,210.0,2520.0,Guntur,Guntur,South)
(D3,Hyderabad,North,Rice 5kg,Grocery,4,280.0,1120.0,Hyderabad,Hyderabad,North)
(D4,Vijayawada,South,Shampoo 200ml,Personal,7,140.0,980.0,Vijayawada,Vijayawada,South)
(D4,Hyderabad,North,Notebook,Stationery,15,40.0,600.0,Hyderabad,Hyderabad,North)
(Rice 5kg,staple)
(Rice 5kg,grain)
(Rice 5kg,bulk)
(Tea 500g,beverage)
(Tea 500g,daily)
(Shampoo 200ml,personal)
(Shampoo 200ml,care)
(Notebook,paper)
(Notebook,school)
$ grep -E 'ERROR|Success|Failed' pig.log | sed -E 's/^[0-9-]+ [0-9:,]+ //' | head -12
Success!
Successfully read 10 records (845 bytes) from: "/user/student/sales/sales.csv"
Successfully stored 3 records (62 bytes) in: "/user/student/out/by_category"
[main] INFO org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.MapReduceLauncher - Success!
Success!
Successfully read 3 records (450 bytes) from: "/user/student/dim/stores.csv"
Successfully read 10 records (845 bytes) from: "/user/student/sales/sales.csv"
Successfully stored 9 records (939 bytes) in: "hdfs://localhost:9000/tmp/temp1251966677/tmp-1038105655"
[main] INFO org.apache.pig.backend.hadoop.executionengine.mapReduceLayer.MapReduceLauncher - Success!
Success!
Successfully read 4 records (471 bytes) from: "/user/student/tags.csv"
Successfully stored 9 records (223 bytes) in: "hdfs://localhost:9000/tmp/temp1251966677/tmp1005279582"
$ hdfs dfs -cat /user/student/out/by_category/part-r-00000
Grocery,4,36,8680.0
Stationery,2,35,1400.0
Personal,1,7,980.0
The Python check, 09_pig_equivalent.py:
OUTPUT
Experiment 9 -- the Pig Latin dataflow, one operator at a time
A = LOAD 'sales' -- 9 rows
store product qty revenue
Vijayawada Rice 5kg 10 2,800
Vijayawada Shampoo 200ml 5 700
Guntur Tea 500g 8 1,680
Vijayawada Rice 5kg 6 1,680
... 5 more
B = FILTER A BY qty >= 6 -- 7 rows
store product qty revenue
Vijayawada Rice 5kg 10 2,800
Guntur Tea 500g 8 1,680
Vijayawada Rice 5kg 6 1,680
Hyderabad Notebook 20 800
... 3 more
C = GROUP B BY category -- 3 groups
(Grocery, {4 tuples})
(Personal, {1 tuples})
(Stationery, {2 tuples})
GROUP in Pig produces a BAG per key, not an aggregate.
The bag is the value, and FOREACH ... GENERATE is what turns
it into numbers. Hive fuses the two; Pig keeps them apart,
which is why Pig can do things to a group that SQL cannot
express without a window function
D = FOREACH C GENERATE group, SUM(B.qty), SUM(B.revenue)
category units revenue orders
Grocery 36 8,680 4
Personal 7 980 1
Stationery 35 1,400 2
E = ORDER D BY revenue DESC
top category: Grocery at 8,680
F = JOIN A BY store_key, stores BY store_key
region store revenue
South Vijayawada 6,160
South Guntur 4,200
North Hyderabad 2,520
the operators, and their SQL equivalents:
Pig Latin SQL
LOAD / STORE no equivalent -- SQL assumes a table exists
FILTER WHERE
FOREACH .. GENERATE SELECT
GROUP GROUP BY, but the bag is kept
JOIN JOIN
ORDER ORDER BY
DISTINCT DISTINCT
FLATTEN UNNEST / LATERAL VIEW explode
ILLUSTRATE no equivalent -- sample data through the plan
two operators have no SQL equivalent, and they are the
reason to reach for Pig: LOAD, which lets a script read a
semi-structured file with no schema declared in advance,
and ILLUSTRATE, which pushes a few representative rows
through every step of the plan so you can see where a
12-stage pipeline went wrong
lazy evaluation, which surprises everyone:
nothing runs until STORE or DUMP.
A = LOAD ...; B = FILTER ...; C = GROUP ...;
-- no job has been submitted yet
STORE C INTO 'out'; -- NOW Pig compiles and runs it
because Pig sees the whole dataflow before executing, it
can merge the FILTER into the LOAD and fuse consecutive
FOREACHes into one MapReduce job. Writing the steps
separately costs nothing -- which is the entire argument
against nesting sub-queries to avoid 'extra passes'
GROUP PRODUCES A BAG, NOT AN AGGREGATE
(Grocery, {4 tuples}) — the bag is the value, and FOREACH … GENERATE turns
it into numbers. Hive fuses the two; Pig keeps them apart, which is why Pig
can do things to a group that SQL cannot express without a window function.
The two operators with no SQL equivalent. LOAD reads a semi-structured file with no
schema declared in advance — SQL assumes a table exists. ILLUSTRATE pushes sample rows
through every step of the plan. Those two are the reason to reach for Pig on ETL.
Lazy evaluation. Nothing runs until STORE or DUMP. Pig sees the whole dataflow first,
so it merges the FILTER into the LOAD and fuses consecutive FOREACHes into one job — the
EXPLAIN above shows the plan it made. Writing the steps separately costs nothing — which
is the whole argument against nesting sub-queries "to avoid extra passes".
The join hint that matters. USING 'replicated' is a map-side join: the small relation
loads into every mapper's memory and there is no shuffle. It dies with an OutOfMemoryError
if the relation does not fit — which is why broadcast joins have a size threshold.
RESULT
Pig, on the cluster, gives Grocery 4 orders, 36 units, ₹8,680; Stationery 2, 35, ₹1,400; Personal 1, 7, ₹980 — the bulk orders by category, revenue first. store is a Pig keyword and cannot name a field.
Execute Hive queries for structured data analysis, with tables and partitions.
Lay a table over the sales file, load a partitioned and bucketed table from it, aggregate by region, prune a partition and rank with a window function; then check each figure in DuckDB.
On the cluster, 10_hive.hql:
The Python check, 10_hive_duckdb.py:
THE CROSS-COURSE CHECK
| Region | Revenue | Profit | Margin |
|---|---|---|---|
| South | 10,360 | 2,760 | 26.64% |
| North | 2,520 | 765 | 30.36% |
South = ₹10,360 is the same number Business Intelligence Tools' DAX CALCULATE measure
produced, and the same one Spark produces in experiment 17. Hive printed it from the cluster;
DuckDB, running the same query text, asserts it. Three engines, two languages, one dataset —
asserted, so drift fails the suite.
On the cluster, 10_hive.hql:
-- Experiment 10 -- Hive queries for structured data analysis -- tables and partitions
--
-- Run it: hive -f 10_hive.hql, with sales.csv in /user/student/sales. It was run on a Hadoop 3.3.6 cluster where these labs
-- are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
-- [Changed: this said the file had never been run, as the Hadoop stack could not be
-- installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
--
-- The runnable half is 10_hive_duckdb.py, which runs the same SQL through DuckDB
--
-- run with: hive -f 10_hive.hql or beeline -u jdbc:hive2://localhost:10000
-- Step 1: Create the database
CREATE DATABASE IF NOT EXISTS retail;
USE retail;
-- Step 2: Lay an external table over the file
-- an EXTERNAL table over data you did not produce: DROP TABLE will not
-- delete the files. Use this for anything you cannot recreate.
CREATE EXTERNAL TABLE IF NOT EXISTS sales_raw (
date_key STRING,
store STRING,
region STRING,
product STRING,
category STRING,
qty INT,
list_price DOUBLE
)
ROW FORMAT DELIMITED FIELDS TERMINATED BY ','
STORED AS TEXTFILE
LOCATION '/user/student/sales/'
TBLPROPERTIES ('skip.header.line.count'='1');
-- Step 3: Create the partitioned, bucketed table
-- the table you actually query: PARTITIONED and columnar
CREATE TABLE IF NOT EXISTS sales (
date_key STRING,
store STRING,
region STRING,
product STRING,
category STRING,
qty INT,
list_price DOUBLE,
revenue DOUBLE
)
PARTITIONED BY (quarter STRING)
CLUSTERED BY (store) INTO 3 BUCKETS
STORED AS PARQUET;
-- Step 4: Load it with dynamic partitions
-- dynamic partitioning: Hive reads the LAST select column as the partition
SET hive.exec.dynamic.partition = true;
SET hive.exec.dynamic.partition.mode = nonstrict;
INSERT OVERWRITE TABLE sales PARTITION (quarter)
SELECT date_key, store, region, product, category, qty, list_price,
qty * list_price AS revenue,
CASE WHEN date_key IN ('D1','D2') THEN 'Q1' ELSE 'Q2' END AS quarter
FROM sales_raw;
SHOW PARTITIONS sales;
DESCRIBE FORMATTED sales;
-- Step 5: Aggregate by region
-- the aggregate verified against Course 11's DAX and experiment 10's DuckDB
SELECT region, SUM(revenue) AS revenue, SUM(qty) AS units
FROM sales
GROUP BY region
ORDER BY revenue DESC;
-- expected: South 10360, North 2520
-- Step 6: Prune a partition
-- PARTITION PRUNING: read one directory, not the table
EXPLAIN DEPENDENCY SELECT SUM(revenue) FROM sales WHERE quarter = 'Q2';
-- input_partitions lists quarter=Q2 alone: one directory is read. If the
-- filter is on a non-partition column it lists every partition, and no
-- amount of indexing will save it -- Hive dropped indexes in version 3.
-- [Corrected: this was a plain EXPLAIN, with the advice to look for
-- "partition values:[Q2]" in the plan. Hive 3's plan does not print that
-- line; its TableScan says 4 rows (Q2's four of nine) but names no partition.
-- EXPLAIN DEPENDENCY names the partitions a query reads.]
SELECT category, product, SUM(qty) AS units, SUM(revenue) AS revenue
FROM sales
WHERE quarter = 'Q2' -- prunes DIRECTORIES, before the job starts
GROUP BY category, product
HAVING SUM(revenue) > 1000 -- filters GROUPS, in the reducer
ORDER BY revenue DESC;
-- Step 7: Rank with a window function
-- a window function, which is where Hive stops looking like MapReduce
SELECT region, product, revenue,
RANK() OVER (PARTITION BY region ORDER BY revenue DESC) AS rnk
FROM sales;
-- Step 8: Tune, and compute statistics
-- the settings worth knowing
-- SET hive.execution.engine = tez; -- mr is deprecated and slow
-- [Corrected: this line ran. Tez is a separate install, and where it is not
-- installed the next query dies with NoClassDefFoundError:
-- org/apache/tez/runtime/api/Event -- and the CLI then hangs instead of
-- exiting. Set it only where Tez is installed; this runs on MapReduce.]
SET hive.vectorized.execution.enabled = true;
SET hive.cbo.enable = true; -- cost-based optimiser, needs stats
ANALYZE TABLE sales PARTITION(quarter) COMPUTE STATISTICS FOR COLUMNS;
The Python check, 10_hive_duckdb.py:
"""Experiment 10 -- Hive queries for structured data analysis: tables,
partitions and the queries that go with them.
`10_hive.hql` carries the HiveQL you submit, and runs on Hive itself (the lab
page shows it). This runs the same questions through DuckDB, which speaks close
enough to ANSI SQL that the SAME query text answers them -- so the figures are
asserted here in seconds, and Hive's answers must agree with them.
The data is Course 11's star schema, imported rather than copied, so a Hive
aggregate here and a DAX measure there are computed from the same nine rows.
"""
import duckdb
import fixtures as f
def q(con, sql):
return con.execute(sql).fetchall()
def main():
print(" Experiment 10 -- Hive-style SQL over the star schema")
con = duckdb.connect()
con.register("sales", f.SALES_DF)
print(f"\n {len(f.SALES_DF)} fact rows, "
f"total revenue {f.total_revenue():,.0f}")
# Step 1: Aggregate by region
rows = q(con, """
SELECT region, SUM(revenue) AS revenue, SUM(profit) AS profit
FROM sales GROUP BY region ORDER BY revenue DESC
""")
print(f"\n {'region':<10}{'revenue':>12}{'profit':>10}{'margin':>9}")
for region, rev, prof in rows:
print(f" {region:<10}{rev:>12,.0f}{prof:>10,.0f}{100 * prof / rev:>8.2f}%")
got = dict((r[0], r[1]) for r in rows)
assert got["South"] == 10360.0, "must match Course 11's CALCULATE figure"
assert got["North"] == 2520.0
assert sum(got.values()) == f.total_revenue()
print(""" South = 10,360 is the SAME number Course 11's DAX
CALCULATE measure produced. Two engines, two languages, one
dataset -- if they ever disagree the suite fails, which is
what makes the cross-check worth having""")
# Step 2: Partition by quarter
print("\n partitioning by quarter -- what Hive actually does:")
parts = q(con, """
SELECT quarter, COUNT(*) AS rows, SUM(revenue) AS revenue
FROM sales GROUP BY quarter ORDER BY quarter
""")
total_rows = sum(p[1] for p in parts)
print(f" {'partition':<14}{'rows':>6}{'revenue':>12}{'scanned for Q2':>16}")
for quarter, n, rev in parts:
print(f" quarter={quarter:<7}{n:>6}{rev:>12,.0f}"
f"{(n if quarter == 'Q2' else 0):>16}")
q2_rows = next(n for quarter, n, _ in parts if quarter == "Q2")
print(f" {'TOTAL':<14}{total_rows:>6}{'':>12}{q2_rows:>16}")
assert total_rows == 9 and q2_rows == 4
print(f""" a partitioned table stores each quarter in its own HDFS
DIRECTORY, so 'WHERE quarter = ''Q2''' reads {q2_rows} rows instead
of {total_rows} -- partition PRUNING, decided before a single byte is
read. The partition column is a directory name, not a column
in the data files, which is why it costs no storage""")
print("""
the trap: partition on something with FEW distinct values.
Partitioning by date_key here would make 4 directories for 9
rows -- the small-files problem from experiment 4, created on
purpose. Partition by quarter or month; BUCKET by customer_id""")
# Step 3: Bucket by store
print("\n bucketing (CLUSTERED BY store INTO 3 BUCKETS):")
buckets = {}
for store in sorted(f.SALES_DF["store"].unique()):
h = sum(ord(c) for c in store) % 3
buckets.setdefault(h, []).append(store)
for b in range(3):
print(f" bucket {b}: {buckets.get(b, []) or '(empty)'}")
assert 1 not in buckets, "bucket 1 draws nothing from three store names"
print(""" BUCKET 1 IS EMPTY, with three stores over three buckets.
Hashing does not distribute small key sets evenly, and an
empty bucket is still a file the job opens.
Buckets are FILES inside a partition, assigned by a hash
of the column. Two tables bucketed the same way on the same
column can be joined bucket-to-bucket with no shuffle at all
-- a sort-merge bucket join, and the reason bucketing exists""")
# Step 4: Compare managed with external tables
print("\n managed against external tables:")
print(f" {'':<12}{'data lives':<26}{'DROP TABLE deletes'}")
print(f" {'MANAGED':<12}{'/user/hive/warehouse':<26}{'the DATA too'}")
print(f" {'EXTERNAL':<12}{'wherever you point it':<26}{'only the metadata'}")
print(""" use EXTERNAL for data you did not produce and cannot
recreate. A DROP TABLE on a managed table over the company's
only copy of a dataset is the classic Hive accident""")
# Step 5: Join, as Hive does
rows = q(con, """
SELECT category, product, SUM(qty) AS units, SUM(revenue) AS revenue
FROM sales GROUP BY category, product
HAVING SUM(revenue) > 1000
ORDER BY revenue DESC
""")
print(f"\n products above 1,000 revenue:")
print(f" {'category':<12}{'product':<16}{'units':>6}{'revenue':>10}")
for cat, prod, units, rev in rows:
print(f" {cat:<12}{prod:<16}{units:>6.0f}{rev:>10,.0f}")
assert len(rows) == 4, "four products clear 1,000 -- Notebook only just"
grocery = sum(r[3] for r in rows if r[0] == "Grocery")
assert grocery == 9800.0
print(""" HAVING filters GROUPS, WHERE filters ROWS -- and in Hive
that distinction is a job-plan difference, not a syntax
nicety: a WHERE on a partition column prunes directories
before the job starts, a HAVING cannot""")
# Step 6: See what Hive is not
print("\n what Hive is NOT:")
print(f" {'expectation':<30}{'reality'}")
for exp, real in (
("row-level UPDATE/DELETE", "only with ACID tables + ORC + buckets"),
("sub-second queries", "seconds to minutes -- it plans a JOB"),
("indexes", "removed in Hive 3; use partitions and ORC/Parquet"),
("a running server holding data", "metadata only; data is files in HDFS"),
("enforced constraints", "declarative only; NOT enforced")):
print(f" {exp:<30}{real}")
print(""" Hive is a COMPILER: HiveQL in, a MapReduce/Tez/Spark job
out. Everything surprising about it follows from that one
sentence, and it is the right answer to 'compare Hive with an
RDBMS'""")
con.close()
if __name__ == "__main__":
main()
On the cluster, 10_hive.hql:
OUTPUT
$ hdfs dfs -mkdir -p /user/student/sales /user/hive/warehouse /tmp/hive
$ hdfs dfs -chmod 733 /tmp/hive && hdfs dfs -chmod g+w /user/hive/warehouse
$ hdfs dfs -put sales.csv /user/student/sales/
$ schematool -dbType derby -initSchema 2>&1 | grep -E 'schemaTool completed'
schemaTool completed
$ hive -f 10_hive.hql 2>hive.log | grep -v '^WARN: '
quarter=Q1
quarter=Q2
# col_name data_type comment
date_key string
store string
region string
product string
category string
qty int
list_price double
revenue double
# Partition Information
# col_name data_type comment
quarter string
# Detailed Table Information
Database: retail
OwnerType: USER
Owner: root
CreateTime: Sun Oct 04 23:21:28 UTC 2026
LastAccessTime: UNKNOWN
Retention: 0
Location: hdfs://localhost:9000/user/hive/warehouse/retail.db/sales
Table Type: MANAGED_TABLE
Table Parameters:
COLUMN_STATS_ACCURATE {\"BASIC_STATS\":\"true\"}
bucketing_version 2
numFiles 6
numPartitions 2
numRows 9
rawDataSize 72
totalSize 6690
transient_lastDdlTime 1791156088
# Storage Information
SerDe Library: org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe
InputFormat: org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat
OutputFormat: org.apache.hadoop.hive.ql.io.parquet.MapredParquetOutputFormat
Compressed: No
Num Buckets: 3
Bucket Columns: [store]
Sort Columns: []
Storage Desc Params:
serialization.format 1
South 10360.0 48
North 2520.0 39
{"input_tables":[{"tablename":"retail@sales","tabletype":"MANAGED_TABLE"}],"input_partitions":[{"partitionName":"retail@sales@quarter=Q2"}]}
Grocery Tea 500g 12 2520.0
Grocery Rice 5kg 4 1120.0
North Rice 5kg 1120.0 1
North Notebook 800.0 2
North Notebook 600.0 3
South Rice 5kg 2800.0 1
South Tea 500g 2520.0 2
South Tea 500g 1680.0 3
South Rice 5kg 1680.0 3
South Shampoo 200ml 980.0 5
South Shampoo 200ml 700.0 6
$ awk '/^FAILED|^Error|Exception/' hive.log | head -5
The Python check, 10_hive_duckdb.py:
OUTPUT
Experiment 10 -- Hive-style SQL over the star schema
9 fact rows, total revenue 12,880
region revenue profit margin
South 10,360 2,760 26.64%
North 2,520 765 30.36%
South = 10,360 is the SAME number Course 11's DAX
CALCULATE measure produced. Two engines, two languages, one
dataset -- if they ever disagree the suite fails, which is
what makes the cross-check worth having
partitioning by quarter -- what Hive actually does:
partition rows revenue scanned for Q2
quarter=Q1 5 7,660 0
quarter=Q2 4 5,220 4
TOTAL 9 4
a partitioned table stores each quarter in its own HDFS
DIRECTORY, so 'WHERE quarter = ''Q2''' reads 4 rows instead
of 9 -- partition PRUNING, decided before a single byte is
read. The partition column is a directory name, not a column
in the data files, which is why it costs no storage
the trap: partition on something with FEW distinct values.
Partitioning by date_key here would make 4 directories for 9
rows -- the small-files problem from experiment 4, created on
purpose. Partition by quarter or month; BUCKET by customer_id
bucketing (CLUSTERED BY store INTO 3 BUCKETS):
bucket 0: ['Guntur', 'Hyderabad']
bucket 1: (empty)
bucket 2: ['Vijayawada']
BUCKET 1 IS EMPTY, with three stores over three buckets.
Hashing does not distribute small key sets evenly, and an
empty bucket is still a file the job opens.
Buckets are FILES inside a partition, assigned by a hash
of the column. Two tables bucketed the same way on the same
column can be joined bucket-to-bucket with no shuffle at all
-- a sort-merge bucket join, and the reason bucketing exists
managed against external tables:
data lives DROP TABLE deletes
MANAGED /user/hive/warehouse the DATA too
EXTERNAL wherever you point it only the metadata
use EXTERNAL for data you did not produce and cannot
recreate. A DROP TABLE on a managed table over the company's
only copy of a dataset is the classic Hive accident
products above 1,000 revenue:
category product units revenue
Grocery Rice 5kg 20 5,600
Grocery Tea 500g 20 4,200
Personal Shampoo 200ml 12 1,680
Stationery Notebook 35 1,400
HAVING filters GROUPS, WHERE filters ROWS -- and in Hive
that distinction is a job-plan difference, not a syntax
nicety: a WHERE on a partition column prunes directories
before the job starts, a HAVING cannot
what Hive is NOT:
expectation reality
row-level UPDATE/DELETE only with ACID tables + ORC + buckets
sub-second queries seconds to minutes -- it plans a JOB
indexes removed in Hive 3; use partitions and ORC/Parquet
a running server holding data metadata only; data is files in HDFS
enforced constraints declarative only; NOT enforced
Hive is a COMPILER: HiveQL in, a MapReduce/Tez/Spark job
out. Everything surprising about it follows from that one
sentence, and it is the right answer to 'compare Hive with an
RDBMS'
PARTITION PRUNING
| Partition | Rows | Revenue | Scanned for Q2 |
|---|---|---|---|
quarter=Q1 |
5 | 7,660 | 0 |
quarter=Q2 |
4 | 5,220 | 4 |
| total | 9 | 12,880 | 4 of 9 |
A partition is an HDFS directory, so WHERE quarter='Q2' reads one
directory — decided before a byte is read; Hive's EXPLAIN DEPENDENCY above lists the one
partition it will read. The partition column is a directory name, so it costs no storage.
The trap: partitioning by date_key here would make 4 directories for 9
rows — the small-files problem, created on purpose.
BUCKETING, AND AN HONEST RESULT
CLUSTERED BY (store) INTO 3 BUCKETS over three stores:
bucket 0: ['Guntur', 'Hyderabad']
bucket 1: (empty)
bucket 2: ['Vijayawada']
Bucket 1 is empty. Hashing does not distribute small key sets evenly, and an empty bucket is still a file the job opens. Reporting that is worth more than pretending the hash was balanced.
Managed against external. DROP TABLE on a MANAGED table deletes the data. Use
EXTERNAL for anything you did not produce and cannot recreate — the classic Hive accident.
HAVING against WHERE. Four products clear ₹1,000 (Grocery total ₹9,800). WHERE
filters rows, HAVING filters groups — and in Hive that is a job-plan difference: a WHERE on
a partition column prunes directories before the job starts, a HAVING cannot.
WHAT HIVE IS NOT
| Expectation | Reality |
|---|---|
row-level UPDATE/DELETE |
only with ACID + ORC + buckets |
| sub-second queries | seconds to minutes — it plans a job |
| indexes | removed in Hive 3 |
| a server holding data | metadata only |
| enforced constraints | declarative, not enforced |
Hive is a compiler. Everything surprising follows from that sentence.
RESULT
Hive, on the cluster, gives South ₹10,360 from 48 units and North ₹2,520 from 39 — the number Business Intelligence Tools' DAX produced — and its query for Q2 reads one partition. DuckDB agrees on every figure.
Import data from an RDBMS into Hadoop using Sqoop.
Import a MySQL table into HDFS in parallel, import a query and a table into Hive, import only the new rows, save the import as a job, and export a summary back; then model the split queries Sqoop runs.
On the cluster, 11_sqoop.sh:
The Python check, 11_sqoop_equivalent.py:
SPLITTING BY THE PRIMARY KEY
step 1: SELECT MIN(order_id), MAX(order_id) -> 1, 90
step 2: four ranges, one per mapper
| Mapper | WHERE |
Rows |
|---|---|---|
| 0 | order_id >= 1 AND <= 22 |
22 |
| 1 | order_id >= 23 AND <= 45 |
23 |
| 2 | order_id >= 46 AND <= 67 |
22 |
| 3 | order_id >= 68 AND <= 90 |
23 |
Four TCP connections to the database. Sqoop's parallelism is database
parallelism — -m 20 against a production OLTP box is a denial of service you
wrote yourself. The Python half has a real SQLite database at one end and a real Parquet file
at the other. The cluster's run prints the first step: BoundingValsQuery: SELECT
MIN(order_id), MAX(order_id) FROM orders.
On the cluster, 11_sqoop.sh:
# Experiment 11 -- import data from an RDBMS into Hadoop using Sqoop
#
# Run it: bash 11_sqoop.sh, with MySQL and a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 11_sqoop_equivalent.py, which does the same import from a real SQLite database
#
# The database: MySQL (MariaDB) on this machine, with retail.orders -- 90 rows,
# ten copies of the nine shared sales rows -- and retail.customers.
# [Corrected: the commands connected to dbhost, a name for your database
# server; it is localhost here. Every command below ran.]
# Step 1: Store the password
echo -n "student-pw" > .pw
hdfs dfs -put .pw /user/student/.pw && rm .pw
hdfs dfs -chmod 400 /user/student/.pw
# [Changed: these three lines are added -- the commands below read the file.]
# Step 2: List the databases and tables
sqoop list-databases --connect jdbc:mysql://localhost:3306 \
--username student --password-file /user/student/.pw 2>/dev/null
sqoop list-tables --connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw 2>/dev/null
# NEVER use --password on the command line: it lands in `ps` and in the
# shell history. --password-file, on HDFS, mode 400.
# Step 3: Import a table, in parallel
# --split-by order_id Sqoop runs SELECT MIN/MAX on THIS column
# --num-mappers 4 4 range queries, 4 connections
# the null flags without them, SQL NULL becomes the literal string
# "null" and every downstream count is wrong
sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders \
--split-by order_id \
--num-mappers 4 \
--target-dir /user/student/orders \
--as-parquetfile \
--compress --compression-codec snappy \
--null-string '\\N' --null-non-string '\\N' \
2>&1 | grep -E "BoundingValsQuery|Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hdfs dfs -ls /user/student/orders | awk 'NR>1{print $NF}' | grep -v "^/user/student/orders/\.metadata"
# [Corrected: the comments sat after the backslashes, as `--split-by order_id \ # ...`.
# A backslash then escapes the space, not the newline, so the command ended at the
# comment and each following option ran as a command of its own -- `--num-mappers:
# command not found`. A comment cannot sit on a continued line.]
# Step 4: Import a query
# $CONDITIONS is MANDATORY and not optional decoration: Sqoop substitutes
# each mapper's range predicate there. Omit it and the import fails.
sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--query 'SELECT o.*, c.region AS cust_region FROM orders o JOIN customers c
ON o.cust_id = c.id WHERE $CONDITIONS' \
--split-by o.order_id \
--target-dir /user/student/orders_enriched \
2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hdfs dfs -cat /user/student/orders_enriched/part-m-00000 | head -2
# [Corrected: the query was `SELECT o.*, c.region`. orders has a region column
# of its own, so the result had two columns named region, and Sqoop, which makes
# one Java field per column, stopped: "Import failed: Duplicate Column identifier
# specified: 'region'". Every column of a --query import needs its own name.]
# Step 5: Import into Hive
export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:$HIVE_HOME/lib/hive-common-3.1.3.jar
# [Corrected: Sqoop reads Hive's settings with Hive's own HiveConf class, and
# without that jar on HADOOP_CLASSPATH it fails with "ClassNotFoundException:
# org.apache.hadoop.hive.conf.HiveConf". Only that jar: with all of Hive's lib/,
# as is often advised, Sqoop runs Hive inside its own JVM, under a security
# manager it installs, and Hive's Derby metastore is refused -- "access denied
# org.apache.derby.security.SystemPermission( "engine", "usederbyinternals" )",
# ten retries, then failure. With just HiveConf it runs the hive command instead.]
sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table customers -m 1 \
--hive-import --hive-database retail --hive-table customers \
--create-hive-table 2>&1 | grep -E "Hive import complete|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
hive -e "SELECT region, COUNT(*) FROM retail.customers GROUP BY region" 2>/dev/null | grep -v "^WARN"
# [Corrected: this was `sqoop import ... --hive-import`, with three dots where
# the connection and the table go; and a table with no primary key needs -m 1.]
# Step 6: Import only what is new
# ten new orders arrive in the database
mysql -h 127.0.0.1 -u student -pstudent-pw retail -e \
"INSERT INTO orders SELECT order_id + 90, cust_id, store, region, product, category, qty, revenue, NOW()
FROM orders WHERE order_id <= 10"
sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders --target-dir /user/student/orders_inc -m 1 \
--incremental append --check-column order_id --last-value 90 \
2>&1 | grep -E "Retrieved [0-9]+ records|--last-value|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
# lastmodified needs a timestamp column, and --merge-key so an updated row
# REPLACES its old copy rather than being added beside it:
# sqoop import ... --incremental lastmodified --check-column updated_at \
# --last-value '2026-08-01 00:00:00' --merge-key order_id
# a saved job REMEMBERS --last-value for you
sqoop job --create orders_inc -- import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders --target-dir /user/student/orders_job -m 1 \
--incremental append --check-column order_id --last-value 0 2>/dev/null
sqoop job --exec orders_inc 2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
sqoop job --show orders_inc 2>/dev/null | grep -E "incremental.last.value|incremental.col"
# [Corrected: these were `sqoop job --create orders_inc -- import ...`, with the
# connection and table left out as three dots.]
# [Noted: Sqoop 1.4.7 does not ship the org.json jar its saved jobs need. Without
# it, --create fails with "NoClassDefFoundError: org/json/JSONObject" yet keeps
# the job's name, with none of its options, and --exec then says "--table or
# --query is required for import". setup_hadoop.sh adds the jar.]
# Step 7: Export back to the database
# an export is NOT transactional across mappers. If mapper 3 fails, the
# rows mappers 1 and 2 wrote are already committed. Use --staging-table
# when that matters.
mysql -h 127.0.0.1 -u student -pstudent-pw retail -B -N -e \
"SELECT category, COUNT(*), SUM(qty), SUM(revenue) FROM orders GROUP BY category" \
| tr '\t' ',' > by_category.csv
hdfs dfs -mkdir -p /user/student/out/by_category
hdfs dfs -put by_category.csv /user/student/out/by_category/
sqoop export \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table order_summary \
--export-dir /user/student/out/by_category \
--update-mode allowinsert --update-key category \
--batch 2>&1 | grep -E "Exported [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
mysql -h 127.0.0.1 -u student -pstudent-pw retail -e "SELECT * FROM order_summary ORDER BY category"
# [Changed: the summary to export is made here from the database itself; the
# script assumed it was already on HDFS.]
# --- the three that bite ----------------------------------------------------
# 1. no primary key and no --split-by -> Sqoop refuses; use -m 1
# 2. --split-by on a skewed column -> one mapper does most of the work
# 3. neither mode notices a DELETE -> re-import fully, periodically
The Python check, 11_sqoop_equivalent.py:
"""Experiment 11 -- import data from an RDBMS into Hadoop using Sqoop.
`11_sqoop.sh` carries the real commands, and runs with Sqoop against MySQL on a
cluster (the lab page shows it). What runs here is the same import in Python:
a REAL relational database (SQLite), a REAL split-by query, REAL parallel range
reads, and a REAL Parquet file at the other end.
Sqoop's entire trick is one line of SQL you never see:
SELECT MIN(id), MAX(id) FROM table
and then one range query per mapper. Everything students find confusing about
Sqoop -- why --split-by matters, why a text primary key breaks it, why 4
mappers can produce wildly unequal files -- follows from that.
"""
import os
import sqlite3
import tempfile
import pyarrow as pa
import pyarrow.parquet as pq
import fixtures as f
FIELDS = ["order_id", "store", "region", "product", "category", "qty", "revenue"]
def build_rdbms(path, rows):
con = sqlite3.connect(path)
con.execute("""CREATE TABLE orders (
order_id INTEGER PRIMARY KEY, store TEXT, region TEXT,
product TEXT, category TEXT, qty INTEGER, revenue REAL)""")
con.executemany("INSERT INTO orders VALUES (?,?,?,?,?,?,?)", rows)
con.commit()
return con
def boundaries(con, table, col, mappers):
"""Exactly what Sqoop runs before it launches a single mapper."""
lo, hi = con.execute(f"SELECT MIN({col}), MAX({col}) FROM {table}").fetchone()
step = (hi - lo + 1) / mappers
out = []
for m in range(mappers):
a = lo + int(m * step)
b = lo + int((m + 1) * step) - 1 if m < mappers - 1 else hi
out.append((a, b))
return lo, hi, out
def main():
print(" Experiment 11 -- Sqoop import, with a real database at one end")
tmp = tempfile.mkdtemp(prefix="bigdata11_")
db = os.path.join(tmp, "retail.db")
# Step 1: Build the source table
# 90 orders, built from the nine shared rows so the totals stay checkable
base = f.SALES_DF
rows = []
for i in range(90):
r = base.iloc[i % 9]
rows.append((i + 1, r["store"], r["region"], r["product"],
r["category"], int(r["qty"]), float(r["revenue"])))
con = build_rdbms(db, rows)
n = con.execute("SELECT COUNT(*) FROM orders").fetchone()[0]
total = con.execute("SELECT SUM(revenue) FROM orders").fetchone()[0]
print(f"\n source: SQLite table 'orders', {n} rows, "
f"revenue {total:,.0f}")
assert n == 90 and abs(total - f.total_revenue() * 10) < 1e-6
print(f""" the 90 rows are ten copies of Course 11's nine, so the
source total is exactly 10 x {f.total_revenue():,.0f}. If the imported
Parquet does not carry that number, the import lost data --
and that is the only import test that matters""")
# Step 2: Find the split boundaries
lo, hi, ranges = boundaries(con, "orders", "order_id", 4)
print(f"\n step 1: SELECT MIN(order_id), MAX(order_id) -> {lo}, {hi}")
print(f" step 2: split into 4 ranges, one per mapper")
print(f" {'mapper':<8}{'WHERE clause':<44}{'rows':>6}")
imported = []
for m, (a, b) in enumerate(ranges):
where = f"order_id >= {a} AND order_id <= {b}"
part = con.execute(
f"SELECT {', '.join(FIELDS)} FROM orders WHERE {where}").fetchall()
imported.extend(part)
print(f" {m:<8}{where:<44}{len(part):>6}")
counts = [len(con.execute(
f"SELECT 1 FROM orders WHERE order_id >= {a} AND order_id <= {b}"
).fetchall()) for a, b in ranges]
assert sum(counts) == n
assert max(counts) - min(counts) <= 1, "an integer key splits near-evenly"
print(f""" four mappers, {min(counts)} or {max(counts)} rows each ({n} does not
divide by 4), four TCP connections to
the database. Sqoop's parallelism is DATABASE parallelism --
raise -m to 20 on a production OLTP box and you have written
a denial of service against your own company""")
# Step 3: Split by a skewed column
print("\n now split by a column that is NOT uniform -- 'qty':")
lo2, hi2, ranges2 = boundaries(con, "orders", "qty", 4)
print(f" MIN(qty), MAX(qty) = {lo2}, {hi2}")
print(f" {'mapper':<8}{'range':<20}{'rows':>6}")
skew = []
for m, (a, b) in enumerate(ranges2):
c = con.execute(
f"SELECT COUNT(*) FROM orders WHERE qty >= {a} AND qty <= {b}"
).fetchone()[0]
skew.append(c)
print(f" {m:<8}{f'{a}..{b}':<20}{c:>6}")
assert sum(skew) == n
assert max(skew) > 3 * min(skew) if min(skew) else True
print(f""" {max(skew)} rows for one mapper and {min(skew)} for another.
Sqoop assumes the split column is UNIFORMLY DISTRIBUTED
between its min and max, and qty is not. The job's wall
clock is the slowest mapper, so a bad --split-by wastes
three quarters of your parallelism.
Split on the PRIMARY KEY unless you have measured otherwise""")
print("\n --split-by on a TEXT column:")
print(" Sqoop needs an ORDERED, NUMERIC column to compute ranges.")
print(" On text it must either refuse, or use")
print(" -Dorg.apache.sqoop.splitter.allow_text_splitter=true")
print(" which splits on string ordering and skews horribly.")
print(""" a table with a UUID or composite primary key has no
natural split column, and the honest answer is -m 1 --
one mapper, no parallelism, correct results""")
# Step 4: Write the target
table = pa.Table.from_pylist(
[dict(zip(FIELDS, r)) for r in sorted(imported)])
out = os.path.join(tmp, "orders.parquet")
pq.write_table(table, out, compression="snappy")
back = pq.read_table(out)
imported_total = sum(back.column("revenue").to_pylist())
print(f"\n landed: {out.split(os.sep)[-1]}, {back.num_rows} rows, "
f"revenue {imported_total:,.0f}")
assert back.num_rows == n
assert abs(imported_total - total) < 1e-6, "the import must not lose money"
print(""" row count AND the sum of a money column, both checked.
Counting rows alone would not catch a truncated numeric
type, which is the classic Sqoop bug: an Oracle NUMBER(38)
silently becoming a Java double""")
# Step 5: Import incrementally
print("\n incremental import, the two modes:")
print(f" {'mode':<15}{'--check-column':<18}{'catches'}")
print(f" {'append':<15}{'an increasing id':<18}"
f"{'new rows only'}")
print(f" {'lastmodified':<15}{'a timestamp':<18}"
f"{'new AND updated rows'}")
last = con.execute("SELECT MAX(order_id) FROM orders").fetchone()[0]
con.executemany("INSERT INTO orders VALUES (?,?,?,?,?,?,?)",
[(91, "Guntur", "South", "Tea 500g", "Grocery", 3, 630.0)])
con.commit()
new = con.execute(
f"SELECT COUNT(*) FROM orders WHERE order_id > {last}").fetchone()[0]
assert new == 1
print(f"\n --last-value {last} now selects {new} row")
print(""" NEITHER MODE CATCHES A DELETE. Sqoop has no way to see a
row that is gone, so an incrementally imported table drifts
away from its source over time. The fix is a periodic full
re-import, and knowing that is the difference between having
used Sqoop and having read about it""")
con.close()
os.remove(out)
os.remove(db)
os.rmdir(tmp)
if __name__ == "__main__":
main()
On the cluster, 11_sqoop.sh:
OUTPUT
$ hdfs dfs -mkdir -p /user/student /user/hive/warehouse /tmp/hive
$ hdfs dfs -chmod 733 /tmp/hive && hdfs dfs -chmod g+w /user/hive/warehouse
$ schematool -dbType derby -initSchema 2>&1 | grep -E 'schemaTool completed'
schemaTool completed
$ hive -e 'CREATE DATABASE retail' 2>/dev/null | grep -v '^WARN'
$ echo -n "student-pw" > .pw
$ hdfs dfs -put .pw /user/student/.pw && rm .pw
$ hdfs dfs -chmod 400 /user/student/.pw
$ sqoop list-databases --connect jdbc:mysql://localhost:3306 \
--username student --password-file /user/student/.pw 2>/dev/null
information_schema
test
retail
$ sqoop list-tables --connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw 2>/dev/null
customers
order_summary
orders
$ sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders \
--split-by order_id \
--num-mappers 4 \
--target-dir /user/student/orders \
--as-parquetfile \
--compress --compression-codec snappy \
--null-string '\\N' --null-non-string '\\N' \
2>&1 | grep -E "BoundingValsQuery|Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO db.DataDrivenDBInputFormat: BoundingValsQuery: SELECT MIN(`order_id`), MAX(`order_id`) FROM `orders`
INFO mapreduce.ImportJobBase: Retrieved 90 records.
$ hdfs dfs -ls /user/student/orders | awk 'NR>1{print $NF}' | grep -v "^/user/student/orders/\.metadata"
/user/student/orders/.signals
/user/student/orders/01ea0149-6607-4381-9487-10a9b64e080e.parquet
/user/student/orders/037af4d2-e6a3-43c3-98a4-c055218b97ec.parquet
/user/student/orders/24ad948c-16db-4ef2-bb18-3ce84fbfd6ad.parquet
/user/student/orders/df60c83e-1143-4106-80b7-ab7fc1e3a2fd.parquet
$ sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--query 'SELECT o.*, c.region AS cust_region FROM orders o JOIN customers c
ON o.cust_id = c.id WHERE $CONDITIONS' \
--split-by o.order_id \
--target-dir /user/student/orders_enriched \
2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 90 records.
$ hdfs dfs -cat /user/student/orders_enriched/part-m-00000 | head -2
1,1,Vijayawada,South,Rice 5kg,Grocery,10,2800.0,2026-08-01 09:00:00.0,South
2,2,Vijayawada,South,Shampoo 200ml,Personal,5,700.0,2026-08-01 09:00:00.0,South
$ export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:$HIVE_HOME/lib/hive-common-3.1.3.jar
$ sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table customers -m 1 \
--hive-import --hive-database retail --hive-table customers \
--create-hive-table 2>&1 | grep -E "Hive import complete|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO hive.HiveImport: Hive import complete.
$ hive -e "SELECT region, COUNT(*) FROM retail.customers GROUP BY region" 2>/dev/null | grep -v "^WARN"
North 2
South 4
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -e \
"INSERT INTO orders SELECT order_id + 90, cust_id, store, region, product, category, qty, revenue, NOW()
FROM orders WHERE order_id <= 10"
$ sqoop import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders --target-dir /user/student/orders_inc -m 1 \
--incremental append --check-column order_id --last-value 90 \
2>&1 | grep -E "Retrieved [0-9]+ records|--last-value|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 10 records.
INFO tool.ImportTool: --last-value 100
$ sqoop job --create orders_inc -- import \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table orders --target-dir /user/student/orders_job -m 1 \
--incremental append --check-column order_id --last-value 0 2>/dev/null
$ sqoop job --exec orders_inc 2>&1 | grep -E "Retrieved [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ImportJobBase: Retrieved 100 records.
$ sqoop job --show orders_inc 2>/dev/null | grep -E "incremental.last.value|incremental.col"
incremental.last.value = 100
incremental.col = order_id
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -B -N -e \
"SELECT category, COUNT(*), SUM(qty), SUM(revenue) FROM orders GROUP BY category" \
| tr '\t' ',' > by_category.csv
$ hdfs dfs -mkdir -p /user/student/out/by_category
$ hdfs dfs -put by_category.csv /user/student/out/by_category/
$ sqoop export \
--connect jdbc:mysql://localhost:3306/retail \
--username student --password-file /user/student/.pw \
--table order_summary \
--export-dir /user/student/out/by_category \
--update-mode allowinsert --update-key category \
--batch 2>&1 | grep -E "Exported [0-9]+ records|ERROR|Exception:" | sed -E 's/^[0-9-]+ [0-9:,]+ //'
INFO mapreduce.ExportJobBase: Exported 3 records.
$ mysql -h 127.0.0.1 -u student -pstudent-pw retail -e "SELECT * FROM order_summary ORDER BY category"
category orders units revenue
Grocery 56 450 110600
Personal 22 132 18480
Stationery 22 385 15400
The Python check, 11_sqoop_equivalent.py:
OUTPUT
Experiment 11 -- Sqoop import, with a real database at one end
source: SQLite table 'orders', 90 rows, revenue 128,800
the 90 rows are ten copies of Course 11's nine, so the
source total is exactly 10 x 12,880. If the imported
Parquet does not carry that number, the import lost data --
and that is the only import test that matters
step 1: SELECT MIN(order_id), MAX(order_id) -> 1, 90
step 2: split into 4 ranges, one per mapper
mapper WHERE clause rows
0 order_id >= 1 AND order_id <= 22 22
1 order_id >= 23 AND order_id <= 45 23
2 order_id >= 46 AND order_id <= 67 22
3 order_id >= 68 AND order_id <= 90 23
four mappers, 22 or 23 rows each (90 does not
divide by 4), four TCP connections to
the database. Sqoop's parallelism is DATABASE parallelism --
raise -m to 20 on a production OLTP box and you have written
a denial of service against your own company
now split by a column that is NOT uniform -- 'qty':
MIN(qty), MAX(qty) = 4, 20
mapper range rows
0 4..7 40
1 8..11 20
2 12..15 20
3 16..20 10
40 rows for one mapper and 10 for another.
Sqoop assumes the split column is UNIFORMLY DISTRIBUTED
between its min and max, and qty is not. The job's wall
clock is the slowest mapper, so a bad --split-by wastes
three quarters of your parallelism.
Split on the PRIMARY KEY unless you have measured otherwise
--split-by on a TEXT column:
Sqoop needs an ORDERED, NUMERIC column to compute ranges.
On text it must either refuse, or use
-Dorg.apache.sqoop.splitter.allow_text_splitter=true
which splits on string ordering and skews horribly.
a table with a UUID or composite primary key has no
natural split column, and the honest answer is -m 1 --
one mapper, no parallelism, correct results
landed: orders.parquet, 90 rows, revenue 128,800
row count AND the sum of a money column, both checked.
Counting rows alone would not catch a truncated numeric
type, which is the classic Sqoop bug: an Oracle NUMBER(38)
silently becoming a Java double
incremental import, the two modes:
mode --check-column catches
append an increasing id new rows only
lastmodified a timestamp new AND updated rows
--last-value 90 now selects 1 row
NEITHER MODE CATCHES A DELETE. Sqoop has no way to see a
row that is gone, so an incrementally imported table drifts
away from its source over time. The fix is a periodic full
re-import, and knowing that is the difference between having
used Sqoop and having read about it
The database is MariaDB, made by _drive_11_sqoop.py: retail.orders with 90 rows — ten
copies of the nine sales rows, ₹128,800, exactly ten times Business Intelligence Tools'
₹12,880 — and retail.customers, with no primary key. Running the script found three faults,
noted in it: comments after the line-continuing backslashes, which ended each command early; the
query's two columns named region; and Hive's jars on the classpath, needed for the Hive import
— but only its HiveConf, or the import fails inside Sqoop's own security manager. And Sqoop
1.4.7 does not ship the org.json jar its saved jobs need; setup_hadoop.sh adds it.
SPLITTING BY A SKEWED COLUMN
| Mapper | qty range |
Rows |
|---|---|---|
| 0 | 4..7 | 40 |
| 1 | 8..11 | 20 |
| 2 | 12..15 | 20 |
| 3 | 16..20 | 10 |
Forty against ten. Sqoop assumes the split column is uniformly distributed
between min and max. The job's wall clock is the slowest mapper, so a bad
--split-by wastes three quarters of your parallelism.
The import is verified two ways. 90 rows and ₹128,800 both check out. Counting rows alone
would not catch a truncated numeric type — an Oracle NUMBER(38) silently becoming a Java
double is the classic Sqoop corruption bug.
NEITHER INCREMENTAL MODE CATCHES A DELETE
--last-value 90 selects the new rows — 10 on the cluster, 1 in the model. But Sqoop has no
way to see a row that is gone, so an incrementally imported table drifts from its source. The
fix is a periodic full re-import — and knowing that is the difference between having used Sqoop
and having read about it.
RESULT
From MariaDB: 90 orders imported by four mappers, 90 rows of a join, the customers table into Hive — South 4, North 2 — 10 new rows incrementally, 100 by a saved job that then remembers 100, and 3 summary rows exported back. In the model, the four range queries split 22, 23, 22, 23, and a skewed column splits 40 against 10.
Capture and store log or streaming data using Flume.
Run a Flume agent that tails a web server's access log, adds headers with interceptors, routes server errors to an alert sink and the rest to HDFS; then model its channel and back-pressure.
On the cluster, 12_flume.conf:
The Python check, 12_flume_equivalent.py:
THE INTERCEPTOR
headers {'host': '10.0.0.1', 'status': '200'}
body (unchanged, 78 chars)
An interceptor adds headers and leaves the body alone. Headers are what a multiplexing selector routes on, so "send 500s to the alert sink" is a header rule, not code.
On the cluster, 12_flume.conf:
# Experiment 12 -- capture and store log/streaming data using Flume
#
# Run it: the flume-ng command below, with a cluster running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 12_flume_equivalent.py, which runs the source/channel/sink semantics
#
# run with:
# flume-ng agent --conf $FLUME_HOME/conf --conf-file 12_flume.conf \
# --name a1 -Dflume.root.logger=INFO,console
# Step 1: Name the agent's parts
a1.sources = tailsrc
a1.channels = memch filech
a1.sinks = hdfssink alertsink
# Step 2: Tail the log files
a1.sources.tailsrc.type = TAILDIR
a1.sources.tailsrc.filegroups = f1
a1.sources.tailsrc.filegroups.f1 = /tmp/flume-lab/logs/access.log.*
a1.sources.tailsrc.positionFile = /tmp/flume-lab/taildir_position.json
# TAILDIR, not `exec tail -F`: exec sources lose everything on restart
# because there is no position file. This is the single most common
# Flume data-loss bug.
# Step 3: Add headers with interceptors
a1.sources.tailsrc.interceptors = ts host regex
a1.sources.tailsrc.interceptors.ts.type = timestamp
a1.sources.tailsrc.interceptors.host.type = host
a1.sources.tailsrc.interceptors.host.useIP = true
a1.sources.tailsrc.interceptors.regex.type = regex_extractor
a1.sources.tailsrc.interceptors.regex.regex = ^\\S+ \\S+ \\S+ \\[[^\\]]+\\] "\\S+ (\\S+)[^"]*" (\\d{3})
a1.sources.tailsrc.interceptors.regex.serializers = s1 s2
a1.sources.tailsrc.interceptors.regex.serializers.s1.name = path
a1.sources.tailsrc.interceptors.regex.serializers.s2.name = status
# [Corrected: the regex was written with single backslashes, ^\S+ \S+ ... (\d{3}).
# Flume reads this file as Java properties, where a backslash escapes the next
# character and is dropped -- \S became S and \d became d -- so the regex matched
# no line, no event got a status header, and every one went to the default
# channel. In a properties file every regex backslash is doubled.]
# Step 4: Route on a header
a1.sources.tailsrc.selector.type = multiplexing
a1.sources.tailsrc.selector.header = status
a1.sources.tailsrc.selector.mapping.500 = filech
a1.sources.tailsrc.selector.default = memch
# Step 5: Set up the channels
a1.channels.memch.type = memory
a1.channels.memch.capacity = 10000
a1.channels.memch.transactionCapacity = 1000
# capacity is EVENTS, not bytes. A full channel blocks the source --
# back-pressure, which is correct behaviour and looks like a hang.
a1.channels.filech.type = file
a1.channels.filech.checkpointDir = /tmp/flume-lab/checkpoint
a1.channels.filech.dataDirs = /tmp/flume-lab/data
# file channel: survives a crash, roughly 10x slower. "We must not lose
# events" and `type = memory` cannot both be true.
# Step 6: Set up the sinks
a1.sinks.hdfssink.type = hdfs
a1.sinks.hdfssink.hdfs.path = hdfs://localhost:9000/logs/dt=%Y-%m-%d/hr=%H
a1.sinks.hdfssink.hdfs.filePrefix = access
a1.sinks.hdfssink.hdfs.fileType = DataStream
a1.sinks.hdfssink.hdfs.writeFormat = Text
a1.sinks.hdfssink.hdfs.batchSize = 1000
# 10 min, not the 30 s default; one HDFS block; 0 = never roll on count
a1.sinks.hdfssink.hdfs.rollInterval = 600
a1.sinks.hdfssink.hdfs.rollSize = 134217728
a1.sinks.hdfssink.hdfs.rollCount = 0
# THE DEFAULTS (30 s / 1 KB / 10 events) MANUFACTURE THE SMALL-FILES
# PROBLEM. Left alone they produce a file every few seconds, each a few
# hundred bytes, and the NameNode pays for every one. Always override.
# [Corrected: the three comments sat after the values, as `rollInterval = 600
# # 10 min`. A properties file has no end-of-line comments: the value was the
# whole rest of the line, the agent refused it -- "NumberFormatException: For
# input string: "600 # 10 min, not the 30 s default"" -- and dropped the
# HDFS sink, so nothing reached HDFS. A comment needs a line of its own.]
a1.sinks.alertsink.type = logger
# [Corrected: the paths were a server's -- /var/log/apache2, /var/lib/flume, and
# a NameNode called nn. Here the agent tails a copy of the lab's access log in
# /tmp/flume-lab, keeps its state there, and writes to the NameNode on localhost.]
# Step 7: Wire it together
a1.sources.tailsrc.channels = memch filech
a1.sinks.hdfssink.channel = memch
a1.sinks.alertsink.channel = filech
The Python check, 12_flume_equivalent.py:
"""Experiment 12 -- capture and store log/streaming data using Flume.
`12_flume.conf` carries the real agent configuration, and runs in Flume itself
(the lab page shows it). What runs here is the AGENT'S SEMANTICS: a source, a
channel with a bounded capacity, a sink with a batch size, and what actually
happens when the sink is slower than the source -- which is the only Flume
question worth asking.
"""
from collections import deque
import fixtures as f
class Channel:
"""A bounded buffer. Flume's memory channel, minus the threads.
The capacity is the whole story: a full channel makes the SOURCE block,
which is back-pressure, which is the correct behaviour and the one that
surprises people.
"""
def __init__(self, capacity):
self.capacity = capacity
self.q = deque()
self.rejected = 0
self.high_water = 0
def put(self, event):
if len(self.q) >= self.capacity:
self.rejected += 1
return False
self.q.append(event)
self.high_water = max(self.high_water, len(self.q))
return True
def take(self, n):
out = [self.q.popleft() for _ in range(min(n, len(self.q)))]
return out
def interceptor(event):
"""Flume interceptors add HEADERS; they do not change the body."""
ip = event.split(" ", 1)[0]
status = event.rsplit(" ", 2)[-2]
return {"headers": {"host": ip, "status": status}, "body": event}
def run_agent(events, capacity, batch, sink_every):
"""Drive one source -> channel -> sink agent, tick by tick."""
chan = Channel(capacity)
delivered, tick = [], 0
src = list(events)
while src or chan.q:
if src:
chan.put(interceptor(src.pop(0)))
if tick % sink_every == 0:
delivered.extend(chan.take(batch))
tick += 1
delivered.extend(chan.take(len(chan.q)))
return chan, delivered, tick
def main():
print(" Experiment 12 -- a Flume agent, semantics first")
logs = f.access_logs(40)
print(f"\n source: {len(logs)} access-log lines")
print(f" {logs[0]}")
# Step 1: Intercept an event
ev = interceptor(logs[0])
print(f"\n after the interceptor:")
print(f" headers {ev['headers']}")
print(f" body (unchanged, {len(ev['body'])} chars)")
assert ev["headers"]["host"] == "10.0.0.1"
assert ev["headers"]["status"] == "200"
assert ev["body"] == logs[0]
print(""" an interceptor adds HEADERS and leaves the body alone.
Headers are what a multiplexing channel selector routes on,
so 'send 500s to the alert sink and everything else to HDFS'
is a header rule, not code""")
# Step 2: Run a healthy agent
chan, delivered, ticks = run_agent(logs, capacity=100, batch=10, sink_every=1)
print(f"\n capacity 100, batch 10, sink every tick:")
print(f" delivered {len(delivered)} of {len(logs)}, "
f"rejected {chan.rejected}, peak channel depth {chan.high_water}")
assert len(delivered) == len(logs) and chan.rejected == 0
assert chan.high_water <= 10
# Step 3: Slow the sink
chan2, delivered2, _ = run_agent(logs, capacity=8, batch=4, sink_every=6)
print(f"\n capacity 8, batch 4, sink every 6th tick (a SLOW sink):")
print(f" delivered {len(delivered2)} of {len(logs)}, "
f"rejected {chan2.rejected}, peak depth {chan2.high_water}")
assert chan2.rejected > 0, "a slow sink must fill the channel"
assert chan2.high_water == 8
print(f""" {chan2.rejected} events were REFUSED by the channel, because the
sink could not drain it. In a real agent the source then
BLOCKS rather than dropping -- back-pressure travels back up
the pipe to the web server. 'Flume lost my events' almost
always means 'the channel was full and the source gave up'""")
# Step 4: Fix it
print("\n the fix, and its cost:")
for cap in (8, 20, 100):
c, d, _ = run_agent(logs, capacity=cap, batch=4, sink_every=6)
print(f" capacity {cap:>4}: rejected {c.rejected:>3}, "
f"peak depth {c.high_water:>3}")
print(""" a bigger channel absorbs a longer burst and buys nothing
if the sink is permanently slower than the source. Buffers
smooth BURSTS; they cannot fix a throughput deficit, and
that sentence answers most Flume tuning questions""")
# Step 5: Compare the channel types
print("\n channel types, and what you are choosing between:")
print(f" {'channel':<12}{'survives a crash?':<20}{'throughput'}")
for name, durable, tput in (
("memory", "NO -- events lost", "highest"),
("file", "yes -- WAL on disk", "roughly 10x slower"),
("Kafka", "yes -- replicated", "high, but another cluster")):
print(f" {name:<12}{durable:<20}{tput}")
print(""" a memory channel plus 'we must not lose events' is a
contradiction, and it is the most common Flume misconfig.
Choose the channel from the durability requirement, then
size the cluster for whatever throughput that leaves""")
# Step 6: Read what the sink writes
from collections import Counter
by_status = Counter(e["headers"]["status"] for e in delivered)
by_host = Counter(e["headers"]["host"] for e in delivered)
print(f"\n what landed in HDFS, by header:")
print(f" status: {dict(sorted(by_status.items()))}")
print(f" host : {dict(sorted(by_host.items()))}")
assert sum(by_status.values()) == 40
assert by_status["200"] == 24 and by_status["404"] == 8
assert len(by_host) == 4 and all(v == 10 for v in by_host.values())
print(""" 24 successes, 8 not-founds and 8 server errors, evenly
over four hosts. That breakdown is what experiment 17 reads
back with Spark -- ingestion and analysis on the same bytes,
which is the point of building the pipeline at all""")
# Step 7: Set the rollover
print("\n the HDFS sink's rollover settings:")
print(f" {'setting':<24}{'default':<12}{'what it does'}")
for k, v, w in (("hdfs.rollInterval", "30 sec", "close the file on a timer"),
("hdfs.rollSize", "1024 bytes", "close it at a size"),
("hdfs.rollCount", "10 events", "close it after N events")):
print(f" {k:<24}{v:<12}{w}")
files = -(-len(logs) // 10)
print(f"\n at the DEFAULTS, {len(logs)} events produce ~{files} HDFS files")
assert files == 4
print(""" and every one of them is a few hundred bytes. Left alone,
Flume's defaults manufacture the small-files problem from
experiment 4 at a rate of two per minute. Set rollCount to 0
and rollSize to a block, or run a compaction job -- this is
the single most common Flume-in-production mistake""")
if __name__ == "__main__":
main()
On the cluster, 12_flume.conf:
OUTPUT
$ rm -rf /tmp/flume-lab && mkdir -p /tmp/flume-lab/logs
$ cp access.log /tmp/flume-lab/logs/access.log.1
$ wc -l < /tmp/flume-lab/logs/access.log.1
40
$ timeout 45 flume-ng agent --conf $FLUME_HOME/conf --conf-file 12_flume.conf --name a1 -Dflume.root.logger=INFO,console > agent.log 2>&1
$ grep -c 'LoggerSink: Event:' agent.log
8
$ grep 'LoggerSink: Event:' agent.log | grep -oE 'path=[^,]+|status=[0-9]+' | paste - - | sort | uniq -c
8 path=/static/app.js status=500
$ hdfs dfs -ls -R /logs | awk '{print $NF}'
/logs/dt=2026-10-04
/logs/dt=2026-10-04/hr=23
/logs/dt=2026-10-04/hr=23/access.1791156576815
$ hdfs dfs -cat '/logs/*/*/*' | wc -l
32
$ hdfs dfs -cat '/logs/*/*/*' | awk '{print $9}' | sort | uniq -c
24 200
8 404
$ awk '/ (ERROR|FATAL) /' agent.log | sed -E 's/^[0-9T:,-]+ //' | sort -u | head -5
The Python check, 12_flume_equivalent.py:
OUTPUT
Experiment 12 -- a Flume agent, semantics first
source: 40 access-log lines
10.0.0.1 - - [12/Aug/2025:09:00:00 +0530] "GET /index.html HTTP/1.1" 200 512
after the interceptor:
headers {'host': '10.0.0.1', 'status': '200'}
body (unchanged, 76 chars)
an interceptor adds HEADERS and leaves the body alone.
Headers are what a multiplexing channel selector routes on,
so 'send 500s to the alert sink and everything else to HDFS'
is a header rule, not code
capacity 100, batch 10, sink every tick:
delivered 40 of 40, rejected 0, peak channel depth 1
capacity 8, batch 4, sink every 6th tick (a SLOW sink):
delivered 32 of 40, rejected 8, peak depth 8
8 events were REFUSED by the channel, because the
sink could not drain it. In a real agent the source then
BLOCKS rather than dropping -- back-pressure travels back up
the pipe to the web server. 'Flume lost my events' almost
always means 'the channel was full and the source gave up'
the fix, and its cost:
capacity 8: rejected 8, peak depth 8
capacity 20: rejected 0, peak depth 16
capacity 100: rejected 0, peak depth 16
a bigger channel absorbs a longer burst and buys nothing
if the sink is permanently slower than the source. Buffers
smooth BURSTS; they cannot fix a throughput deficit, and
that sentence answers most Flume tuning questions
channel types, and what you are choosing between:
channel survives a crash? throughput
memory NO -- events lost highest
file yes -- WAL on disk roughly 10x slower
Kafka yes -- replicated high, but another cluster
a memory channel plus 'we must not lose events' is a
contradiction, and it is the most common Flume misconfig.
Choose the channel from the durability requirement, then
size the cluster for whatever throughput that leaves
what landed in HDFS, by header:
status: {'200': 24, '404': 8, '500': 8}
host : {'10.0.0.1': 10, '10.0.0.2': 10, '10.0.0.3': 10, '10.0.0.4': 10}
24 successes, 8 not-founds and 8 server errors, evenly
over four hosts. That breakdown is what experiment 17 reads
back with Spark -- ingestion and analysis on the same bytes,
which is the point of building the pipeline at all
the HDFS sink's rollover settings:
setting default what it does
hdfs.rollInterval 30 sec close the file on a timer
hdfs.rollSize 1024 bytes close it at a size
hdfs.rollCount 10 events close it after N events
at the DEFAULTS, 40 events produce ~4 HDFS files
and every one of them is a few hundred bytes. Left alone,
Flume's defaults manufacture the small-files problem from
experiment 4 at a rate of two per minute. Set rollCount to 0
and rollSize to a block, or run a compaction job -- this is
the single most common Flume-in-production mistake
_drive_12_flume.py puts the lab's 40 access-log lines where the agent looks, runs it for 45
seconds, and counts what reached each sink. Running the configuration found two faults, both
from its being a Java properties file:
No comment can follow a value. rollInterval = 600 # 10 min made the value the whole
rest of the line; the agent refused it — NumberFormatException: For input string: "600 …" —
and dropped the HDFS sink.
A backslash escapes the next character, and is dropped. The regex ^\S+ … (\d{3}) became
^S+ … (d{3}), matched nothing, and every event went to the default channel. Every regex
backslash is doubled.
THE CHANNEL, AND BACK-PRESSURE
| Configuration | Delivered | Rejected | Peak depth |
|---|---|---|---|
| capacity 100, batch 10, fast sink | 40 | 0 | ≤10 |
| capacity 8, batch 4, slow sink | 40 | 8 | 8 |
| capacity 20, slow sink | 40 | 0 | 16 |
| capacity 100, slow sink | 40 | 0 | 16 |
Eight events refused because the sink could not drain the channel. In a real agent the source then blocks — back-pressure travelling back to the web server. "Flume lost my events" almost always means "the channel was full and the source gave up". And a bigger channel absorbs a longer burst and buys nothing once the sink is permanently slower. Buffers smooth bursts; they cannot fix a throughput deficit.
What landed. {'200': 24, '404': 8, '500': 8}, evenly over four hosts (10 each) — the
agent on the cluster split them the same way. Those same numbers appear in experiments 14 and
17 — three code paths, one set of figures.
THE DEFAULTS MANUFACTURE THE SMALL-FILES PROBLEM
rollInterval 30 s, rollSize 1024 bytes, rollCount 10 events → 40 events
produce ~4 HDFS files, each a few hundred bytes. Left alone, Flume generates
the Unit 2 small-files problem at two files a minute. With the configuration's settings the 32
events made one file.
RESULT
The agent read the 40 log lines, sent the 8 with status 500 to the logger sink and the other 32 — 24 of 200, 8 of 404 — to one HDFS file under dt=/hr=. Before two corrections it delivered nothing to HDFS and routed nothing.
Serialize and store datasets in Avro and Parquet formats.
Write the sales in Avro and in Parquet, read them back, evolve the Avro schema, read only some Parquet columns, and compare the sizes honestly.
The Python check, 13_avro_parquet.py:
THE REAL FORMATS
fastavro and pyarrow are real implementations, so the files written here
are byte-for-byte readable by Hadoop, Hive and Spark.
Avro: self-describing. 9 records, 938 bytes, round-trip exact. The writer's schema is in
the file header — full name in.ac.datascience.sales.Sale, namespace included.
The Python check, 13_avro_parquet.py:
"""Experiment 13 -- serialize and store datasets in Avro and Parquet.
THIS EXPERIMENT FULLY RUNS. fastavro and pyarrow are real implementations of
the real formats, so the files written here are byte-for-byte readable by
Hadoop, Hive and Spark. Nothing is simulated.
The point of the experiment is not "how do I call the library" -- it is the
difference between a ROW format and a COLUMN format, which decides everything
about how a big-data query performs.
"""
import io
import json
import os
import tempfile
import fastavro
import pyarrow as pa
import pyarrow.parquet as pq
import fixtures as f
AVRO_SCHEMA = {
"type": "record",
"name": "Sale",
"namespace": "in.ac.datascience.sales",
"fields": [
{"name": "date_key", "type": "string"},
{"name": "store", "type": "string"},
{"name": "region", "type": "string"},
{"name": "product", "type": "string"},
{"name": "category", "type": "string"},
{"name": "qty", "type": "long"},
{"name": "revenue", "type": "double"},
{"name": "profit", "type": "double"},
],
}
FIELDS = [fld["name"] for fld in AVRO_SCHEMA["fields"]]
def records():
return [{k: (int(r[k]) if k == "qty" else r[k]) for k in FIELDS}
for _, r in f.SALES_DF.iterrows()]
def main():
print(" Experiment 13 -- Avro and Parquet, both really written")
rows = records()
tmp = tempfile.mkdtemp(prefix="bigdata13_")
# Step 1: Write and read Avro
avro_path = os.path.join(tmp, "sales.avro")
with open(avro_path, "wb") as fh:
fastavro.writer(fh, fastavro.parse_schema(AVRO_SCHEMA), rows)
with open(avro_path, "rb") as fh:
back = list(fastavro.reader(fh))
assert back == rows, "Avro must round-trip exactly"
avro_size = os.path.getsize(avro_path)
# The notes quote this figure, so it is asserted. It is deterministic --
# fixed rows, fixed schema -- but NOT independent of the schema's text:
# the namespace is stored in the file header, so renaming it moves the
# byte count. That is how this assertion earns its place; the figure had
# already drifted once, silently, when the namespace changed.
assert avro_size == 938, (
f"Avro file is {avro_size} bytes, the notes say 938 -- "
"update lab.md and unit-4.md, or find out what changed")
print(f"\n Avro : {len(rows)} records, {avro_size} bytes, round-trip exact")
# the schema travels INSIDE the file -- this is the property that matters
with open(avro_path, "rb") as fh:
embedded = fastavro.reader(fh).writer_schema
assert embedded["name"] == "in.ac.datascience.sales.Sale", (
"Avro stores the FULL name -- namespace + name -- not the short one")
assert [fl["name"] for fl in embedded["fields"]] == FIELDS
print(f"\n the embedded schema's full name is {embedded['name']!r}")
print(""" the WRITER'S SCHEMA is stored in the file header, so an
Avro file is self-describing. A reader five years later needs
no external metadata, which is exactly what a CSV cannot
promise -- and why Avro is the ingestion format""")
# Step 2: Evolve the schema
evolved = json.loads(json.dumps(AVRO_SCHEMA))
evolved["fields"].append(
{"name": "channel", "type": ["null", "string"], "default": None})
buf = io.BytesIO()
with open(avro_path, "rb") as fh:
old_bytes = fh.read()
read_new = list(fastavro.reader(io.BytesIO(old_bytes),
reader_schema=fastavro.parse_schema(evolved)))
assert all(r["channel"] is None for r in read_new)
assert len(read_new) == len(rows)
print(f"\n schema evolution: read {len(rows)} OLD records with a NEW schema")
print(f" the added field 'channel' comes back as "
f"{read_new[0]['channel']!r} -- its DEFAULT")
print(""" the old file was NOT rewritten. Avro resolves the writer's
schema against the reader's, field by field, and fills in
defaults for anything missing. A field added WITHOUT a
default breaks exactly this, which is the one rule to
remember about evolving an Avro schema""")
# Step 3: Write Parquet, and read only some columns
table = pa.Table.from_pylist(rows)
pq_path = os.path.join(tmp, "sales.parquet")
pq.write_table(table, pq_path, compression="snappy")
pq_size = os.path.getsize(pq_path)
back_pq = pq.read_table(pq_path).to_pylist()
assert back_pq == rows, "Parquet must round-trip exactly"
print(f"\n Parquet: {len(rows)} records, {pq_size} bytes, round-trip exact")
# column projection -- the whole reason Parquet exists
one_col = pq.read_table(pq_path, columns=["revenue"])
assert one_col.num_columns == 1 and one_col.num_rows == len(rows)
meta = pq.ParquetFile(pq_path).metadata
rg = meta.row_group(0)
col_sizes = {rg.column(i).path_in_schema:
rg.column(i).total_compressed_size
for i in range(rg.num_columns)}
print(f"\n bytes stored PER COLUMN inside the Parquet file:")
for name, sz in sorted(col_sizes.items(), key=lambda kv: -kv[1]):
print(f" {name:<12}{sz:>7}")
total = sum(col_sizes.values())
rev = col_sizes["revenue"]
print(f" {'TOTAL':<12}{total:>7}")
print(f"\n SELECT revenue reads {rev} of {total} column bytes "
f"({100 * rev / total:.1f}%)")
assert rev < total / 4
print(""" THAT is column projection, and it is why Parquet wins on
analytical queries: a SELECT of one column out of eight
reads roughly one column's worth of bytes. A row format has
to read every row in full and discard seven fields""")
# predicate pushdown via row-group statistics
stats = rg.column([i for i in range(rg.num_columns)
if rg.column(i).path_in_schema == "revenue"][0]).statistics
print(f"\n row-group statistics for 'revenue': "
f"min {stats.min:,.0f}, max {stats.max:,.0f}")
assert stats.min == 600.0 and stats.max == 2800.0
print(""" a query for revenue > 5000 can SKIP THIS ENTIRE ROW GROUP
without decoding a byte, because the max is 2,800. That is
predicate pushdown, and on a partitioned Parquet dataset it
is often a bigger win than the compression""")
# Step 4: Compare the two formats
csv_path = os.path.join(tmp, "sales.csv")
f.SALES_DF[FIELDS].to_csv(csv_path, index=False)
csv_size = os.path.getsize(csv_path)
print(f"\n {'format':<12}{'bytes':>8} {'layout':<8}{'schema':<15}{'best for'}")
for name, size, layout, schema, use in (
("CSV", csv_size, "row", "none", "interchange, and nothing else"),
("Avro", avro_size, "row", "in the file", "ingestion, streaming, evolution"),
("Parquet", pq_size, "COLUMN", "in the footer", "analytics, column projection"),
("SequenceFile", None, "row", "external", "legacy Hadoop key/value")):
shown = f"{size:>8}" if size else f"{'--':>8}"
print(f" {name:<12}{shown} {layout:<8}{schema:<15}{use}")
print(f"""
on NINE ROWS Parquet is LARGER than CSV ({pq_size} against
{csv_size}) -- the footer, the schema and the per-column
metadata are fixed overhead that nine rows cannot amortise.
Report that honestly: Parquet's advantage is asymptotic, and
quoting a compression ratio from a toy file is how people
get caught out in a viva""")
# Step 5: Test the claim at size
# 108,000 records. TWO versions: one that repeats the nine rows exactly,
# and one where every row differs -- because a columnar format's headline
# ratio is mostly a statement about how repetitive the data is, and
# quoting the repetitive number alone would be misleading.
big = rows * 12000
varied = [dict(r, qty=r["qty"] + i % 97,
revenue=r["revenue"] + (i % 8191) * 0.25,
store=f"{r['store']}-{i % 500}")
for i, r in enumerate(big)]
vsizes = {}
for name, data in (("repetitive", big), ("varied", varied)):
va = os.path.join(tmp, f"v_{name}.avro")
vp = os.path.join(tmp, f"v_{name}.parquet")
vc = os.path.join(tmp, f"v_{name}.csv")
with open(va, "wb") as fh:
fastavro.writer(fh, fastavro.parse_schema(AVRO_SCHEMA), data)
pq.write_table(pa.Table.from_pylist(data), vp, compression="snappy")
with open(vc, "w") as fh:
fh.write(",".join(FIELDS) + "\n")
for r in data:
fh.write(",".join(str(r[k]) for k in FIELDS) + "\n")
vsizes[name] = {"CSV": os.path.getsize(vc), "Avro": os.path.getsize(va),
"Parquet": os.path.getsize(vp)}
for pth in (va, vp, vc):
os.remove(pth)
print(f"\n the same schema at {len(big):,} records:")
print(f" {'data':<12}{'CSV':>12}{'Avro':>12}{'Parquet':>12}"
f"{'CSV/Parquet':>13}")
for name in ("repetitive", "varied"):
z = vsizes[name]
print(f" {name:<12}{z['CSV']:>12,}{z['Avro']:>12,}"
f"{z['Parquet']:>12,}{z['CSV'] / z['Parquet']:>12.1f}x")
rep = vsizes["repetitive"]["CSV"] / vsizes["repetitive"]["Parquet"]
var = vsizes["varied"]["CSV"] / vsizes["varied"]["Parquet"]
assert rep > var * 5, "the repetitive figure must be visibly inflated"
assert var > 1.5, "Parquet should still beat CSV on varied data"
print(f""" READ BOTH ROWS. The {rep:.0f}x on repetitive data is an
ARTEFACT: 12,000 identical copies of nine rows dictionary-
encode to almost nothing. Give every row a distinct store
and revenue and the ratio falls to {var:.1f}x. Even that is
optimistic -- date, region and category are still repetitive
here -- and Parquet-against-CSV in production usually lands
between 3x and 10x.
A columnar format's headline compression number is mostly a
statement about how REPETITIVE your data is, and a benchmark
on duplicated rows says nothing at all""")
print("\n which format for which job:")
print(" row-by-row WRITES, whole-record reads -> Avro")
print(" column aggregates over billions of rows -> Parquet")
print(" a landing zone that must survive schema")
print(" changes for years -> Avro")
print(" the table Hive and Spark actually query -> Parquet")
print(""" the standard architecture uses BOTH: Avro at the edge
where records arrive one at a time and schemas drift, then a
batch job converts to Parquet for the query layer. That
answer is worth full marks on 'compare Avro and Parquet'""")
for path in (avro_path, pq_path, csv_path):
os.remove(path)
os.rmdir(tmp)
if __name__ == "__main__":
main()
The Python check, 13_avro_parquet.py:
OUTPUT
Experiment 13 -- Avro and Parquet, both really written
Avro : 9 records, 938 bytes, round-trip exact
the embedded schema's full name is 'in.ac.datascience.sales.Sale'
the WRITER'S SCHEMA is stored in the file header, so an
Avro file is self-describing. A reader five years later needs
no external metadata, which is exactly what a CSV cannot
promise -- and why Avro is the ingestion format
schema evolution: read 9 OLD records with a NEW schema
the added field 'channel' comes back as None -- its DEFAULT
the old file was NOT rewritten. Avro resolves the writer's
schema against the reader's, field by field, and fills in
defaults for anything missing. A field added WITHOUT a
default breaks exactly this, which is the one rule to
remember about evolving an Avro schema
Parquet: 9 records, 2584 bytes, round-trip exact
bytes stored PER COLUMN inside the Parquet file:
profit 150
qty 144
revenue 143
product 125
category 111
store 110
date_key 83
region 83
TOTAL 949
SELECT revenue reads 143 of 949 column bytes (15.1%)
THAT is column projection, and it is why Parquet wins on
analytical queries: a SELECT of one column out of eight
reads roughly one column's worth of bytes. A row format has
to read every row in full and discard seven fields
row-group statistics for 'revenue': min 600, max 2,800
a query for revenue > 5000 can SKIP THIS ENTIRE ROW GROUP
without decoding a byte, because the max is 2,800. That is
predicate pushdown, and on a partitioned Parquet dataset it
is often a bigger win than the compression
format bytes layout schema best for
CSV 533 row none interchange, and nothing else
Avro 938 row in the file ingestion, streaming, evolution
Parquet 2584 COLUMN in the footer analytics, column projection
SequenceFile -- row external legacy Hadoop key/value
on NINE ROWS Parquet is LARGER than CSV (2584 against
533) -- the footer, the schema and the per-column
metadata are fixed overhead that nine rows cannot amortise.
Report that honestly: Parquet's advantage is asymptotic, and
quoting a compression ratio from a toy file is how people
get caught out in a viva
the same schema at 108,000 records:
data CSV Avro Parquet CSV/Parquet
repetitive 5,700,058 5,924,195 18,790 303.4x
varied 6,269,519 6,380,511 522,264 12.0x
READ BOTH ROWS. The 303x on repetitive data is an
ARTEFACT: 12,000 identical copies of nine rows dictionary-
encode to almost nothing. Give every row a distinct store
and revenue and the ratio falls to 12.0x. Even that is
optimistic -- date, region and category are still repetitive
here -- and Parquet-against-CSV in production usually lands
between 3x and 10x.
A columnar format's headline compression number is mostly a
statement about how REPETITIVE your data is, and a benchmark
on duplicated rows says nothing at all
which format for which job:
row-by-row WRITES, whole-record reads -> Avro
column aggregates over billions of rows -> Parquet
a landing zone that must survive schema
changes for years -> Avro
the table Hive and Spark actually query -> Parquet
the standard architecture uses BOTH: Avro at the edge
where records arrive one at a time and schemas drift, then a
batch job converts to Parquet for the query layer. That
answer is worth full marks on 'compare Avro and Parquet'
SCHEMA EVOLUTION, DEMONSTRATED
read 9 OLD records with a NEW nine-field schema
the added field 'channel' comes back as None -- its DEFAULT
The old file was not rewritten. Avro resolves writer's schema against reader's, field by field. A field added without a default breaks exactly this, and that is the one rule to remember.
PARQUET: COLUMN PROJECTION
| Column | Bytes |
|---|---|
| profit | 150 |
| qty | 144 |
| revenue | 143 |
| product | 125 |
| category | 111 |
| store | 110 |
| date_key | 83 |
| region | 83 |
| total | 949 |
SELECT revenue reads 143 of 949 column bytes — 15.1%.
Predicate pushdown: row-group statistics for revenue are min 600, max
2,800, so a query for revenue > 5000 skips the entire row group without
decoding a byte.
THE COMPRESSION CLAIM, TOLD HONESTLY
| Data | CSV | Avro | Parquet | CSV/Parquet |
|---|---|---|---|---|
| 9 rows | 533 | 938 | 2,584 | 0.2× |
| 108,000 rows, repetitive | 5,700,058 | 5,924,190 | 18,790 | 303.4× |
| 108,000 rows, varied | 6,269,519 | 6,380,506 | 522,264 | 12.0× |
On nine rows Parquet is 4.8× LARGER than CSV. The 303× is an artefact of 12,000 identical copies. Give every row a distinct store and revenue and it falls to 12.0× — and even that is optimistic. In production, 3× to 10×.
A columnar format's headline ratio is mostly a statement about how repetitive your data is, and a benchmark on duplicated rows says nothing at all.
RESULT
938 bytes of Avro and an exact round trip; old records read with a new schema get the added field's default; SELECT revenue reads 143 of 949 Parquet column bytes. On nine rows Parquet is 4.8 times larger than CSV, and the 303× ratio holds only for repetitive data — 12.0× once the rows vary.
Build an end-to-end ingestion workflow combining batch (Sqoop) and streaming (Flume).
Import the orders in a batch leg and the log events in a stream leg, both to Parquet, then join them in DuckDB — first wrongly, then at a common grain.
The Python check, 14_pipeline.py:
THE FAN TRAP
SQLite → Parquet (batch), log events → Parquet (streaming), then a real DuckDB query across both. The batch leg is experiment 11's import and the stream leg experiment 12's agent, each done here in Python so the two legs can be joined in one program; experiments 11 and 12 ran the real tools on the cluster.
events counted through the join: 90
events actually ingested : 40
Each host appears in several orders, so every event is counted once per matching order. This is the same fan trap Business Intelligence Tools found in a Power BI model — not a SQL problem, a grain problem, appearing wherever two fact tables are joined directly.
The Python check, 14_pipeline.py:
"""Experiment 14 -- an end-to-end ingestion workflow combining batch (Sqoop)
and streaming (Flume).
This is the experiment that ties the course together, and the one where the
interesting problem is not any single tool but the JOIN BETWEEN THEM: batch
data arrives hourly and complete, streaming data arrives continuously and
incomplete, and a query that spans both has to decide what it means by "now".
Runs end to end: SQLite -> Parquet (the batch side), log events -> Parquet
(the streaming side), then a real DuckDB query across the two.
"""
import os
import sqlite3
import tempfile
from collections import Counter
import duckdb
import pyarrow as pa
import pyarrow.parquet as pq
import fixtures as f
def batch_leg(tmp):
"""Sqoop's half: a full-fidelity import of a slow-changing table."""
db = os.path.join(tmp, "orders.db")
con = sqlite3.connect(db)
con.execute("""CREATE TABLE orders (
order_id INTEGER PRIMARY KEY, host TEXT, region TEXT, revenue REAL)""")
rows = []
hosts = ["10.0.0.1", "10.0.0.2", "10.0.0.3", "10.0.0.4"]
for i, (_, r) in enumerate(f.SALES_DF.iterrows()):
rows.append((i + 1, hosts[i % 4], r["region"], float(r["revenue"])))
con.executemany("INSERT INTO orders VALUES (?,?,?,?)", rows)
con.commit()
data = con.execute("SELECT order_id, host, region, revenue FROM orders").fetchall()
con.close()
path = os.path.join(tmp, "batch.parquet")
pq.write_table(pa.Table.from_pylist(
[dict(zip(("order_id", "host", "region", "revenue"), r)) for r in data]),
path)
return path, len(data), sum(r[3] for r in data)
def stream_leg(tmp, n):
"""Flume's half: events, parsed, with headers, landed as Parquet."""
events = []
for line in f.access_logs(n):
host = line.split(" ", 1)[0]
status = line.rsplit(" ", 2)[-2]
size = int(line.rsplit(" ", 1)[-1])
events.append({"host": host, "status": status, "bytes": size})
path = os.path.join(tmp, "stream.parquet")
pq.write_table(pa.Table.from_pylist(events), path)
return path, len(events)
def main():
print(" Experiment 14 -- batch and streaming, joined")
# Step 1: Run the batch leg and the stream leg
tmp = tempfile.mkdtemp(prefix="bigdata14_")
batch, n_batch, batch_rev = batch_leg(tmp)
stream, n_stream = stream_leg(tmp, 40)
print(f"\n batch leg (Sqoop) : {n_batch} orders, "
f"revenue {batch_rev:,.0f}")
print(f" stream leg (Flume) : {n_stream} events")
assert abs(batch_rev - f.total_revenue()) < 1e-6
con = duckdb.connect()
con.execute(f"CREATE VIEW batch AS SELECT * FROM '{batch}'")
con.execute(f"CREATE VIEW stream AS SELECT * FROM '{stream}'")
# Step 2: Join them
rows = con.execute("""
SELECT b.host,
COUNT(DISTINCT b.order_id) AS orders,
SUM(DISTINCT b.revenue) AS revenue,
COUNT(s.host) AS events,
SUM(CASE WHEN s.status = '500' THEN 1 ELSE 0 END) AS errors
FROM batch b LEFT JOIN stream s ON b.host = s.host
GROUP BY b.host ORDER BY b.host
""").fetchall()
print(f"\n joined on host:")
print(f" {'host':<12}{'orders':>8}{'events':>8}{'errors':>8}")
for host, orders, rev, events, errors in rows:
print(f" {host:<12}{orders:>8}{events:>8}{errors:>8}")
total_events = sum(r[3] for r in rows)
assert total_events != n_stream, "the join FANS OUT -- see below"
print(f"""
events counted through the join: {total_events}
events actually ingested : {n_stream}""")
print(""" THE JOIN INFLATED THE EVENT COUNT. Each host appears in
several orders, so every event is counted once per matching
order -- a FAN TRAP, and the same defect Course 11 found in
a Power BI model. It is not a Spark problem, a Hive problem
or a SQL problem; it is a GRAIN problem, and it appears
wherever two fact tables are joined directly""")
# Step 3: Fix the join
fixed = con.execute("""
WITH ev AS (
SELECT host, COUNT(*) AS events,
SUM(CASE WHEN status = '500' THEN 1 ELSE 0 END) AS errors
FROM stream GROUP BY host),
ord AS (
SELECT host, COUNT(*) AS orders, SUM(revenue) AS revenue
FROM batch GROUP BY host)
SELECT o.host, o.orders, o.revenue, e.events, e.errors
FROM ord o JOIN ev e ON o.host = e.host ORDER BY o.host
""").fetchall()
print(f"\n the fix -- aggregate EACH SIDE to a common grain FIRST:")
print(f" {'host':<12}{'orders':>8}{'revenue':>10}{'events':>8}{'errors':>8}")
for host, orders, rev, events, errors in fixed:
print(f" {host:<12}{orders:>8}{rev:>10,.0f}{events:>8}{errors:>8}")
assert sum(r[3] for r in fixed) == n_stream
assert abs(sum(r[2] for r in fixed) - batch_rev) < 1e-6
print(f""" {sum(r[3] for r in fixed)} events and {sum(r[2] for r in fixed):,.0f} revenue -- both totals now
reconcile with the sources. Aggregate to a shared grain, THEN
join. That single rule prevents most wrong numbers in a data
warehouse, and it is worth stating in exactly those words""")
# Step 4: Compare the two legs
print("\n the two legs are not interchangeable:")
print(f" {'':<20}{'batch (Sqoop)':<26}{'streaming (Flume)'}")
for label, b, s in (
("arrives", "on a schedule", "continuously"),
("completeness", "a whole table, consistent", "whatever has landed"),
("late data", "impossible", "NORMAL -- and must be handled"),
("re-runnable", "yes, idempotent", "no -- events are consumed"),
("catches DELETEs", "on a full re-import", "never"),
("file sizes", "large, controllable", "small unless you roll"),
("failure means", "re-run the import", "gap in the data")):
print(f" {label:<20}{b:<26}{s}")
print(""" 'late data is normal' is the row that changes the design.
A streaming aggregate for 09:00 is not final at 10:00, so
either you accept eventual correctness or you keep a
watermark and re-emit. The batch leg has no such problem,
which is why the LAMBDA ARCHITECTURE keeps both""")
# Step 5: Set lambda against kappa
print("\n the two architectures this experiment is really about:")
print(f" {'':<10}{'layers':<34}{'cost'}")
print(f" {'Lambda':<10}{'batch + speed + serving':<34}"
f"{'the logic is written TWICE'}")
print(f" {'Kappa':<10}{'one streaming path, replayable':<34}"
f"{'needs a log like Kafka'}")
print(""" Lambda's real cost is not machines, it is that the same
business rule exists in two codebases and they drift. Kappa
removes the batch layer by making the stream replayable --
which is why Kafka replaced Flume in most of these pipelines
after about 2016. Say that and you have placed the whole
syllabus in time""")
# Step 6: Reconcile
status_counts = Counter(
r[0] for r in con.execute("SELECT status FROM stream").fetchall())
print(f"\n reconciliation: {dict(sorted(status_counts.items()))}")
assert status_counts["200"] == 24
assert sum(status_counts.values()) == 40
print(""" the same 24 / 8 / 8 as experiments 12 and 17. Three
experiments, three code paths, one set of numbers -- and if
a change ever breaks one of them the suite fails on all
three, which is the only reason to build the check""")
con.close()
for path in (batch, stream, os.path.join(tmp, "orders.db")):
os.remove(path)
os.rmdir(tmp)
if __name__ == "__main__":
main()
The Python check, 14_pipeline.py:
OUTPUT
Experiment 14 -- batch and streaming, joined
batch leg (Sqoop) : 9 orders, revenue 12,880
stream leg (Flume) : 40 events
joined on host:
host orders events errors
10.0.0.1 3 30 6
10.0.0.2 2 20 4
10.0.0.3 2 20 4
10.0.0.4 2 20 4
events counted through the join: 90
events actually ingested : 40
THE JOIN INFLATED THE EVENT COUNT. Each host appears in
several orders, so every event is counted once per matching
order -- a FAN TRAP, and the same defect Course 11 found in
a Power BI model. It is not a Spark problem, a Hive problem
or a SQL problem; it is a GRAIN problem, and it appears
wherever two fact tables are joined directly
the fix -- aggregate EACH SIDE to a common grain FIRST:
host orders revenue events errors
10.0.0.1 3 4,200 10 2
10.0.0.2 2 3,220 10 2
10.0.0.3 2 2,800 10 2
10.0.0.4 2 2,660 10 2
40 events and 12,880 revenue -- both totals now
reconcile with the sources. Aggregate to a shared grain, THEN
join. That single rule prevents most wrong numbers in a data
warehouse, and it is worth stating in exactly those words
the two legs are not interchangeable:
batch (Sqoop) streaming (Flume)
arrives on a schedule continuously
completeness a whole table, consistent whatever has landed
late data impossible NORMAL -- and must be handled
re-runnable yes, idempotent no -- events are consumed
catches DELETEs on a full re-import never
file sizes large, controllable small unless you roll
failure means re-run the import gap in the data
'late data is normal' is the row that changes the design.
A streaming aggregate for 09:00 is not final at 10:00, so
either you accept eventual correctness or you keep a
watermark and re-emit. The batch leg has no such problem,
which is why the LAMBDA ARCHITECTURE keeps both
the two architectures this experiment is really about:
layers cost
Lambda batch + speed + serving the logic is written TWICE
Kappa one streaming path, replayable needs a log like Kafka
Lambda's real cost is not machines, it is that the same
business rule exists in two codebases and they drift. Kappa
removes the batch layer by making the stream replayable --
which is why Kafka replaced Flume in most of these pipelines
after about 2016. Say that and you have placed the whole
syllabus in time
reconciliation: {'200': 24, '404': 8, '500': 8}
the same 24 / 8 / 8 as experiments 12 and 17. Three
experiments, three code paths, one set of numbers -- and if
a change ever breaks one of them the suite fails on all
three, which is the only reason to build the check
THE FIX
| host | orders | revenue | events | errors |
|---|---|---|---|---|
| 10.0.0.1 | 3 | 4,200 | 10 | 2 |
| 10.0.0.2 | 2 | 3,220 | 10 | 2 |
| 10.0.0.3 | 2 | 2,800 | 10 | 2 |
| 10.0.0.4 | 2 | 2,660 | 10 | 2 |
| total | 9 | 12,880 | 40 | 8 |
Aggregate each side to a common grain first, then join. Both totals now reconcile with the sources. That single rule prevents most wrong numbers in a data warehouse.
The row that changes the design. "Late data is normal." A streaming aggregate for 09:00 is not final at 10:00, so either you accept eventual correctness or you keep a watermark and re-emit. The batch leg has no such problem — which is exactly why the Lambda architecture keeps both, and why Kappa removes the batch layer by making the stream replayable.
Lambda's real cost is not machines — it is the same business rule living in two codebases and drifting.
RESULT
Joined directly, the 40 events are counted 90 times — the fan trap. Aggregated to the host first, then joined, both sides reconcile: 9 orders, ₹12,880, 40 events, 8 errors.
Create and manage tables in HBase, with CRUD operations.
Create a table with two column families, put, get and scan cells, update and delete them, count atomically, alter and truncate the table and pre-split another; then model row keys, versions and tombstones.
On the cluster, 15_hbase.rb:
The Python check, 15_hbase_model.py:
THE ROW KEY THAT SILENTLY LOSES A SALE
row key 'region#store#date' over 9 fact rows
produces only 8 DISTINCT KEYS -- 1 row would be overwritten
Vijayawada sold Rice and Shampoo on D1. HBase would not complain — it would version one over the other. A row key must be unique at the grain, and in HBase nothing checks that for you.
On the cluster, 15_hbase.rb:
# Experiment 15 -- create and manage tables in HBase -- CRUD operations
#
# Run it: hbase shell 15_hbase.rb, with HBase running. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 15_hbase_model.py, which implements the same data model and runs it
#
# run with: hbase shell 15_hbase.rb (or paste into an interactive shell)
# Step 1: Create the table
create 'sales', \
{NAME => 'info', VERSIONS => 3, COMPRESSION => 'SNAPPY'}, \
{NAME => 'sales', VERSIONS => 3, TTL => 31536000}
# COLUMN FAMILIES ARE FIXED AT CREATE TIME and expensive to change.
# COLUMNS inside a family are free and need no declaration -- that is what
# "schema-less" means in HBase, and it is only half true.
# Keep families to two or three: each one is a separate store file, and a
# flush of one flushes all of them.
list
describe 'sales'
# Step 2: Put cells
# row key = region#store#date#product -- composite, unique at the grain,
# and NOT monotonic. See the design table at the end.
put 'sales', 'South#Vijayawada#D1#Rice', 'info:product', 'Rice 5kg'
put 'sales', 'South#Vijayawada#D1#Rice', 'info:category', 'Grocery'
put 'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '10'
put 'sales', 'South#Vijayawada#D1#Rice', 'sales:revenue', '2800'
put 'sales', 'North#Hyderabad#D2#Notebook', 'info:product', 'Notebook'
put 'sales', 'North#Hyderabad#D2#Notebook', 'sales:qty', '20'
# Step 3: Get a row
get 'sales', 'South#Vijayawada#D1#Rice'
get 'sales', 'South#Vijayawada#D1#Rice', 'sales'
get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
# a PUT to an existing cell ADDS A VERSION; it does not overwrite.
# Step 4: Scan
scan 'sales'
scan 'sales', {LIMIT => 5}
scan 'sales', {STARTROW => 'South', STOPROW => 'South~'}
# '~' sorts after every printable ASCII letter, which is the idiomatic way
# to write a prefix scan by hand. PrefixFilter does the same thing:
scan 'sales', {FILTER => "PrefixFilter('South')"}
scan 'sales', {FILTER => "SingleColumnValueFilter('info','category',=,'binary:Grocery',true,true)"}
# [Corrected: without the two trues -- filterIfMissing and latestVersionOnly --
# a row that has NO info:category passes the filter, and the scan returned the
# Notebook row too. A filter on a column says nothing about rows without it.]
# ^ THIS IS A FULL TABLE SCAN. HBase has no secondary index. The filter runs
# server-side, so less data crosses the network -- but every row is read.
# Step 5: Update and delete
put 'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '11' # = a new version
get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
# [Changed: this get is added. The one above ran before any second put, so it
# showed one version; here are both, 11 over 10.]
delete 'sales', 'South#Vijayawada#D1#Rice', 'info:category'
deleteall 'sales', 'North#Hyderabad#D2#Notebook'
# A DELETE WRITES A TOMBSTONE. The table gets BIGGER. Data and marker are
# both removed only at a major compaction:
major_compact 'sales'
# Step 6: Count atomically
incr 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views', 1
get_counter 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views'
# Step 7: Alter, and truncate
count 'sales', INTERVAL => 100
disable 'sales'
alter 'sales', {NAME => 'info', VERSIONS => 5}
enable 'sales'
truncate 'sales' # = disable + drop + recreate. Keeps the schema.
# drop 'sales' # must be disabled first
# Step 8: Pre-split a table
create 'sales2', 'info', {SPLITS => ['East', 'North', 'South', 'West']}
# without pre-splits the table starts as ONE region on ONE RegionServer, so
# a bulk load runs single-threaded until the first split.
# --- row key design ---------------------------------------------------------
# key writes go to verdict
# timestamp the last region HOTSPOT
# sequential id the last region HOTSPOT
# md5(id) + id everywhere good; range scans lost
# region#store#date#product by region good; prefix scans work
#
# You cannot have even write distribution AND range scans on the same
# dimension. Choosing between them IS row-key design.
The Python check, 15_hbase_model.py:
"""Experiment 15 -- create and manage tables in HBase (CRUD operations).
`15_hbase.rb` carries the real shell commands, and runs in the HBase shell (the
lab page shows it). What runs here is HBase's DATA MODEL, implemented honestly:
a sorted map from (row, family:qualifier, version) to bytes, with real
versioning, real tombstones and real row-key range scans.
The model IS the exam. Almost every HBase question -- why scans are fast and
gets by value are not, why a monotonic row key is a disaster, why a delete
does not free space -- is a consequence of "sorted map, sharded by row-key
range".
"""
import bisect
import fixtures as f
class HBase:
"""(row, family, qualifier) -> {timestamp: value}, kept SORTED BY ROW."""
def __init__(self, families, max_versions=3):
self.families = set(families)
self.max_versions = max_versions
self.cells = {}
self.rows = [] # sorted, because everything depends on it
self.clock = 0
def _tick(self):
self.clock += 1
return self.clock
def put(self, row, fam, qual, value):
if fam not in self.families:
raise KeyError(f"column family {fam!r} was not declared at create time")
if row not in self.cells:
bisect.insort(self.rows, row)
self.cells[row] = {}
versions = self.cells[row].setdefault((fam, qual), {})
versions[self._tick()] = value
for ts in sorted(versions)[:-self.max_versions]:
del versions[ts]
def get(self, row, fam=None, qual=None, versions=1):
if row not in self.cells:
return {}
out = {}
for (fm, q), vs in self.cells[row].items():
if fam and fm != fam:
continue
if qual and q != qual:
continue
live = []
for ts, v in sorted(vs.items(), reverse=True):
if v is None:
break # a tombstone MASKS every older version
live.append((ts, v))
live = live[:versions]
if live:
out[f"{fm}:{q}"] = live if versions > 1 else live[0][1]
return out
def delete(self, row, fam, qual):
"""A delete writes a TOMBSTONE. It does not remove anything."""
self.cells[row].setdefault((fam, qual), {})[self._tick()] = None
def scan(self, start=None, stop=None):
lo = bisect.bisect_left(self.rows, start) if start else 0
hi = bisect.bisect_left(self.rows, stop) if stop else len(self.rows)
return [(r, self.get(r)) for r in self.rows[lo:hi]]
def storefiles(self):
"""Every version and every tombstone still occupies a cell."""
return sum(len(v) for row in self.cells.values() for v in row.values())
def main():
print(" Experiment 15 -- the HBase data model, implemented")
t = HBase(families={"info", "sales"}, max_versions=3)
# Step 1: Create the table and put rows
# First, a row key that looks reasonable and is NOT unique at the grain.
naive = {f"{r['region']}#{r['store']}#{r['date_key']}"
for _, r in f.SALES_DF.iterrows()}
print(f"\n row key 'region#store#date' over {len(f.SALES_DF)} fact rows")
print(f" produces only {len(naive)} DISTINCT KEYS -- "
f"{len(f.SALES_DF) - len(naive)} row would be overwritten")
assert len(naive) == 8, "two facts share a store and a date"
print(""" Vijayawada sold Rice AND Shampoo on D1, so those two
facts collide. HBase would not complain -- it would simply
version one over the other and lose a sale.
A row key must be UNIQUE AT THE GRAIN. In an RDBMS the
primary key declaration catches this; in HBase nothing does,
and that is the failure mode to remember""")
for _, r in f.SALES_DF.iterrows():
# row key: region#store#date#product -- unique, composite, NOT monotonic
key = (f"{r['region']}#{r['store']}#{r['date_key']}#"
f"{r['product'].split()[0]}")
t.put(key, "info", "product", r["product"])
t.put(key, "info", "category", r["category"])
t.put(key, "sales", "qty", int(r["qty"]))
t.put(key, "sales", "revenue", float(r["revenue"]))
print(f"\n {len(t.rows)} rows, {t.storefiles()} cells")
print(f" row keys are SORTED, always:")
for k in t.rows[:4]:
print(f" {k}")
print(f" ... {len(t.rows) - 4} more")
assert t.rows == sorted(t.rows)
# Step 2: Get a row
key = t.rows[0]
print(f"\n GET '{key}':")
for col, val in sorted(t.get(key).items()):
print(f" {col:<18}{val}")
# Step 3: Keep versions
t.put(key, "sales", "qty", 99)
t.put(key, "sales", "qty", 111)
vs = t.get(key, "sales", "qty", versions=3)["sales:qty"]
print(f"\n after two more PUTs to the same cell, 3 versions:")
for ts, v in vs:
print(f" ts={ts:<5}{v}")
assert [v for _, v in vs][:2] == [111, 99]
assert len(vs) == 3, "VERSIONS => 3 caps the history at three"
print(""" a PUT to an existing cell does not overwrite -- it adds a
VERSION, and the old value is still readable. VERSIONS => 3
at create time is what caps it. That is why HBase is
described as a multidimensional map: row, family, qualifier
AND time""")
# Step 4: Delete, with a tombstone
before = t.storefiles()
t.delete(key, "info", "category")
after = t.storefiles()
assert "info:category" not in t.get(key)
assert after > before, "a delete makes the table BIGGER until compaction"
print(f"\n DELETE info:category")
print(f" readable? {'yes' if 'info:category' in t.get(key) else 'no'}")
print(f" cells: {before} -> {after}")
print(""" THE TABLE GOT BIGGER. A delete writes a tombstone marker;
the data and the marker both disappear only at MAJOR
COMPACTION. This is the answer to 'I deleted a billion rows
and disk usage went up'""")
# Step 5: Scan by prefix
south = t.scan("South", "South~")
north = t.scan("North", "North~")
print(f"\n SCAN 'South' .. 'South~' -> {len(south)} rows")
print(f" SCAN 'North' .. 'North~' -> {len(north)} rows")
assert len(south) + len(north) == len(t.rows)
assert len(south) == 6 and len(north) == 3
assert len(t.rows) == 9, "the unique key keeps all nine facts"
print(""" a range scan on the row-key PREFIX reads exactly the
rows you want, sequentially, from one or two regions. That
is the fastest thing HBase does -- and it works only because
'region' is the FIRST component of the key""")
print("\n the same question asked the wrong way:")
matches = [r for r, cols in t.scan() if cols.get("info:category") == "Grocery"]
print(f" find category = 'Grocery' -> {len(matches)} rows, "
f"after scanning all {len(t.rows)}")
print(""" HBase has NO SECONDARY INDEX. Filtering on a value means
a FULL TABLE SCAN with a server-side filter -- correct, and
O(table). If you need that query, you build a second table
keyed by category, and you keep it in sync yourself""")
# Step 6: Design the row key
print("\n row key design, which is the whole job:")
print(f" {'key':<34}{'regions hit by a write':<24}verdict")
for key_desc, hits, verdict in (
("timestamp (1723459200, ...)", "ONE -- always the last", "HOTSPOT"),
("sequential id (1, 2, 3, ...)", "ONE -- always the last", "HOTSPOT"),
("md5(id) + id", "all, evenly", "good, scans lost"),
("region#store#date", "by region", "good, prefix scans work")):
print(f" {key_desc:<34}{hits:<24}{verdict}")
print(""" a monotonically increasing row key sends EVERY write to
the same RegionServer, so a 50-node cluster runs at the
speed of one node. Salting or hashing fixes the hotspot and
destroys range scans -- you cannot have both, and choosing
is what row-key design means""")
# Step 7: Compare HBase with what it is confused with
print("\n HBase against what students compare it to:")
print(f" {'':<14}{'HBase':<26}{'Hive':<22}{'MongoDB (Course 10)'}")
for label, hb, hv, mg in (
("model", "sparse sorted map", "tables over files", "documents"),
("latency", "milliseconds", "seconds to minutes", "milliseconds"),
("random writes", "YES", "no", "YES"),
("secondary index", "no", "no", "YES"),
("query language", "get/put/scan only", "HiveQL", "MQL"),
("schema", "families fixed, cols free", "fixed", "free")):
print(f" {label:<14}{hb:<26}{hv:<22}{mg}")
print(""" HBase and Hive both sit on HDFS and answer completely
different questions: Hive scans everything slowly, HBase
fetches one row instantly. And note the row students always
get wrong -- HBase has NO secondary index where MongoDB
does, which is the sharpest difference between the two
NoSQL stores this programme teaches""")
if __name__ == "__main__":
main()
On the cluster, 15_hbase.rb:
OUTPUT
$ start-hbase.sh > /dev/null 2>&1
$ sleep 20
$ hbase shell -n < 15_hbase.typed 2>&1
hbase:001:0> create 'sales', \
hbase:002:0* {NAME => 'info', VERSIONS => 3, COMPRESSION => 'SNAPPY'}, \
hbase:003:0* {NAME => 'sales', VERSIONS => 3, TTL => 31536000}
Created table sales
Took 1.2161 seconds
=> Hbase::Table - sales
hbase:004:0> list
TABLE
sales
1 row(s)
Took 0.0266 seconds
=> ["sales"]
hbase:005:0> describe 'sales'
Table sales is ENABLED
sales, {TABLE_ATTRIBUTES => {METADATA => {'hbase.store.file-tracker.impl' => 'DEFAULT'}}}
COLUMN FAMILIES DESCRIPTION
{NAME => 'info', INDEX_BLOCK_ENCODING => 'NONE', VERSIONS => '3', KEEP_DELETED_CELLS => 'FALSE', DATA_BLOCK_ENCODING => 'NONE', TTL => 'FOREVER', MIN_VERSIONS => '0', REPLICATION_SCOPE => '0', BLOOMFILTER => 'ROW', IN_MEMORY => 'false', COMPRESSION => 'SNAPPY', BLOCKCACHE => 'true', BLOCKSIZE => '65536 B (64KB)'}
{NAME => 'sales', INDEX_BLOCK_ENCODING => 'NONE', VERSIONS => '3', KEEP_DELETED_CELLS => 'FALSE', DATA_BLOCK_ENCODING => 'NONE', TTL => '31536000 SECONDS (365 DAYS)', MIN_VERSIONS => '0', REPLICATION_SCOPE => '0', BLOOMFILTER => 'ROW', IN_MEMORY => 'false', COMPRESSION => 'NONE', BLOCKCACHE => 'true', BLOCKSIZE => '65536 B (64KB)'}
2 row(s)
Quota is disabled
Took 0.1691 seconds
hbase:006:0> put 'sales', 'South#Vijayawada#D1#Rice', 'info:product', 'Rice 5kg'
Took 0.0882 seconds
hbase:007:0> put 'sales', 'South#Vijayawada#D1#Rice', 'info:category', 'Grocery'
Took 0.0052 seconds
hbase:008:0> put 'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '10'
Took 0.0083 seconds
hbase:009:0> put 'sales', 'South#Vijayawada#D1#Rice', 'sales:revenue', '2800'
Took 0.0065 seconds
hbase:010:0> put 'sales', 'North#Hyderabad#D2#Notebook', 'info:product', 'Notebook'
Took 0.0051 seconds
hbase:011:0> put 'sales', 'North#Hyderabad#D2#Notebook', 'sales:qty', '20'
Took 0.0060 seconds
hbase:012:0> get 'sales', 'South#Vijayawada#D1#Rice'
COLUMN CELL
info:category timestamp=2026-10-04T23:30:58.482, value=Grocery
info:product timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
sales:qty timestamp=2026-10-04T23:30:58.496, value=10
sales:revenue timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0716 seconds
hbase:013:0> get 'sales', 'South#Vijayawada#D1#Rice', 'sales'
COLUMN CELL
sales:qty timestamp=2026-10-04T23:30:58.496, value=10
sales:revenue timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0102 seconds
hbase:014:0> get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
COLUMN CELL
sales:qty timestamp=2026-10-04T23:30:58.496, value=10
1 row(s)
Took 0.0093 seconds
hbase:015:0> scan 'sales'
ROW COLUMN+CELL
North#Hyderabad#D2#Notebook column=info:product, timestamp=2026-10-04T23:30:58.520, value=Notebook
North#Hyderabad#D2#Notebook column=sales:qty, timestamp=2026-10-04T23:30:58.532, value=20
South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
2 row(s)
Took 0.0168 seconds
hbase:016:0> scan 'sales', {LIMIT => 5}
ROW COLUMN+CELL
North#Hyderabad#D2#Notebook column=info:product, timestamp=2026-10-04T23:30:58.520, value=Notebook
North#Hyderabad#D2#Notebook column=sales:qty, timestamp=2026-10-04T23:30:58.532, value=20
South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
2 row(s)
Took 0.0306 seconds
hbase:017:0> scan 'sales', {STARTROW => 'South', STOPROW => 'South~'}
ROW COLUMN+CELL
South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0142 seconds
hbase:018:0> scan 'sales', {FILTER => "PrefixFilter('South')"}
ROW COLUMN+CELL
South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0252 seconds
hbase:019:0> scan 'sales', {FILTER => "SingleColumnValueFilter('info','category',=,'binary:Grocery',true,true)"}
ROW COLUMN+CELL
South#Vijayawada#D1#Rice column=info:category, timestamp=2026-10-04T23:30:58.482, value=Grocery
South#Vijayawada#D1#Rice column=info:product, timestamp=2026-10-04T23:30:58.460, value=Rice 5kg
South#Vijayawada#D1#Rice column=sales:qty, timestamp=2026-10-04T23:30:58.496, value=10
South#Vijayawada#D1#Rice column=sales:revenue, timestamp=2026-10-04T23:30:58.508, value=2800
1 row(s)
Took 0.0670 seconds
hbase:020:0> put 'sales', 'South#Vijayawada#D1#Rice', 'sales:qty', '11' # = a new version
Took 0.0092 seconds
hbase:021:0> get 'sales', 'South#Vijayawada#D1#Rice', {COLUMN => 'sales:qty', VERSIONS => 3}
COLUMN CELL
sales:qty timestamp=2026-10-04T23:30:58.881, value=11
sales:qty timestamp=2026-10-04T23:30:58.496, value=10
1 row(s)
Took 0.0083 seconds
hbase:022:0> delete 'sales', 'South#Vijayawada#D1#Rice', 'info:category'
Took 0.0119 seconds
hbase:023:0> deleteall 'sales', 'North#Hyderabad#D2#Notebook'
Took 0.0048 seconds
hbase:024:0> major_compact 'sales'
Took 0.0647 seconds
hbase:025:0> incr 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views', 1
COUNTER VALUE = 1
Took 0.0173 seconds
hbase:026:0> get_counter 'sales', 'North#Hyderabad#D2#Notebook', 'sales:views'
COUNTER VALUE = 1
Took 0.0045 seconds
hbase:027:0> count 'sales', INTERVAL => 100
2 row(s)
Took 0.0157 seconds
=> 2
hbase:028:0> disable 'sales'
Took 0.7025 seconds
hbase:029:0> alter 'sales', {NAME => 'info', VERSIONS => 5}
Updating all regions with the new schema...
All regions updated.
Done.
Took 1.2158 seconds
hbase:030:0> enable 'sales'
Took 0.6759 seconds
hbase:031:0> truncate 'sales' # = disable + drop + recreate. Keeps the schema.
Truncating 'sales' table (it may take a while):
Disabling table...
Truncating table...
Took 0.9969 seconds
hbase:032:0> create 'sales2', 'info', {SPLITS => ['East', 'North', 'South', 'West']}
Created table sales2
Took 0.6325 seconds
=> Hbase::Table - sales2
hbase:033:0>
$ stop-hbase.sh 2>&1 | tail -1
stopping hbase..............
The Python check, 15_hbase_model.py:
OUTPUT
Experiment 15 -- the HBase data model, implemented
row key 'region#store#date' over 9 fact rows
produces only 8 DISTINCT KEYS -- 1 row would be overwritten
Vijayawada sold Rice AND Shampoo on D1, so those two
facts collide. HBase would not complain -- it would simply
version one over the other and lose a sale.
A row key must be UNIQUE AT THE GRAIN. In an RDBMS the
primary key declaration catches this; in HBase nothing does,
and that is the failure mode to remember
9 rows, 36 cells
row keys are SORTED, always:
North#Hyderabad#D2#Notebook
North#Hyderabad#D3#Rice
North#Hyderabad#D4#Notebook
South#Guntur#D1#Tea
... 5 more
GET 'North#Hyderabad#D2#Notebook':
info:category Stationery
info:product Notebook
sales:qty 20
sales:revenue 800.0
after two more PUTs to the same cell, 3 versions:
ts=38 111
ts=37 99
ts=19 20
a PUT to an existing cell does not overwrite -- it adds a
VERSION, and the old value is still readable. VERSIONS => 3
at create time is what caps it. That is why HBase is
described as a multidimensional map: row, family, qualifier
AND time
DELETE info:category
readable? no
cells: 38 -> 39
THE TABLE GOT BIGGER. A delete writes a tombstone marker;
the data and the marker both disappear only at MAJOR
COMPACTION. This is the answer to 'I deleted a billion rows
and disk usage went up'
SCAN 'South' .. 'South~' -> 6 rows
SCAN 'North' .. 'North~' -> 3 rows
a range scan on the row-key PREFIX reads exactly the
rows you want, sequentially, from one or two regions. That
is the fastest thing HBase does -- and it works only because
'region' is the FIRST component of the key
the same question asked the wrong way:
find category = 'Grocery' -> 5 rows, after scanning all 9
HBase has NO SECONDARY INDEX. Filtering on a value means
a FULL TABLE SCAN with a server-side filter -- correct, and
O(table). If you need that query, you build a second table
keyed by category, and you keep it in sync yourself
row key design, which is the whole job:
key regions hit by a write verdict
timestamp (1723459200, ...) ONE -- always the last HOTSPOT
sequential id (1, 2, 3, ...) ONE -- always the last HOTSPOT
md5(id) + id all, evenly good, scans lost
region#store#date by region good, prefix scans work
a monotonically increasing row key sends EVERY write to
the same RegionServer, so a 50-node cluster runs at the
speed of one node. Salting or hashing fixes the hotspot and
destroys range scans -- you cannot have both, and choosing
is what row-key design means
HBase against what students compare it to:
HBase Hive MongoDB (Course 10)
model sparse sorted map tables over files documents
latency milliseconds seconds to minutes milliseconds
random writes YES no YES
secondary indexno no YES
query languageget/put/scan only HiveQL MQL
schema families fixed, cols free fixed free
HBase and Hive both sit on HDFS and answer completely
different questions: Hive scans everything slowly, HBase
fetches one row instantly. And note the row students always
get wrong -- HBase has NO secondary index where MongoDB
does, which is the sharpest difference between the two
NoSQL stores this programme teaches
_drive_15_hbase.py starts HBase in standalone mode — its own ZooKeeper, data in a local
folder — with the Java Snappy codec configured, since the table asks for COMPRESSION =>
'SNAPPY' and the native library is not there; then types the file into hbase shell.
Versions. A put to an existing cell adds a version; it does not overwrite. The shell's
get with VERSIONS => 3 above shows both values of sales:qty, newest first; the model:
ts=38 111
ts=37 99
ts=19 20
A DELETE MAKES THE TABLE BIGGER
DELETE info:category readable? no cells: 38 -> 39
A tombstone. Data and marker disappear only at major compaction — the answer to "I deleted a billion rows and disk usage went up".
Scans. SCAN 'South'..'South~' → 6 rows in the model's nine, and 'North'..'North~' → 3;
in the shell's two-row table, 1. A prefix
scan reads exactly the rows you want, sequentially — and works only because region is the
first component of the key. Ask it the wrong way — category = 'Grocery' — and it is a
full table scan. HBase has no secondary index.
ROW KEY DESIGN
| Key | Writes go to | Verdict |
|---|---|---|
| timestamp | the last region | HOTSPOT |
| sequential id | the last region | HOTSPOT |
md5(id) + id |
everywhere | good — scans lost |
region#store#date#product |
by region | good — prefix scans work |
You cannot have even write distribution and range scans on the same dimension. Choosing is what row-key design means.
RESULT
In the HBase shell every operation ran: a second put kept both versions of the cell, a prefix scan and a value filter each found the one South row, the counter counted to 1. The filter needed two more arguments, or rows without the column passed it. In the model, the row key region#store#date loses a sale.
Demonstrate coordination with ZooKeeper.
Start a three-server ensemble, see which server leads, build a tree of znodes, make ephemeral and sequential nodes, and run the leader-election recipe; then model election, locking and ensemble sizing.
On the cluster, 16_zookeeper.sh:
The Python check, 16_zookeeper_model.py:
LEADER ELECTION
nn1 -> lock-0000000000 LEADER: nn1
nn2 -> lock-0000000001
nn3 -> lock-0000000002
nn1's session expires: new LEADER: nn2
Nobody ran a failover script. The ephemeral node was deleted by the server, the watch fired, and nn2 saw itself at the head of the queue. That is how HDFS NameNode HA works — the link back to experiment 5.
On the cluster, 16_zookeeper.sh:
# Experiment 16 -- demonstrate coordination with ZooKeeper
#
# Run it: bash 16_zookeeper.sh -- it starts and stops its own ensemble. It was run on a Hadoop 3.3.6 cluster where these labs
# are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
# [Changed: this said the file had never been run, as the Hadoop stack could not be
# installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
#
# The runnable half is 16_zookeeper_model.py, which runs leader election, locking and quorum maths
#
# Step 1: Configure three servers
# conf/zoo.cfg, identical on all three:
# tickTime=2000
# initLimit=10
# syncLimit=5
# dataDir=/var/lib/zookeeper
# clientPort=2181
# server.1=zk1:2888:3888
# server.2=zk2:2888:3888
# server.3=zk3:2888:3888
# and on each host: echo <id> > /var/lib/zookeeper/myid
#
# THREE, FIVE OR SEVEN. Four servers need a quorum of 3 and tolerate one
# failure -- exactly what three tolerate -- so the fourth machine buys nothing
# and slows every write.
#
# Here the three servers run on ONE machine, so each needs its own client port,
# its own pair of quorum ports and its own data directory; otherwise the file
# is the one above. The four-letter commands must also be allowed by name:
for id in 1 2 3; do
mkdir -p zk$id && echo $id > zk$id/myid
cat > zk$id.cfg <<EOF
tickTime=2000
initLimit=10
syncLimit=5
dataDir=$PWD/zk$id
clientPort=218$id
server.1=localhost:2888:3888
server.2=localhost:2889:3889
server.3=localhost:2890:3890
4lw.commands.whitelist=srvr,mntr
EOF
done
# [Corrected: the ensemble was described in comments, and `zkServer.sh start`
# started ONE server. These are the zoo.cfg above for three servers on one
# host, and each is started below.]
# Step 2: Start them, and see the leader
for id in 1 2 3; do ZOO_LOG_DIR=zk$id zkServer.sh start $PWD/zk$id.cfg 2>&1 | tail -1; done
sleep 10 # they elect a leader
for id in 1 2 3; do zkServer.sh status $PWD/zk$id.cfg 2>&1 | grep Mode; done
# "Mode: leader" on exactly one server
echo srvr | nc localhost 2181 | grep -E "Mode|Zxid|Node count" # state and zxid
echo mntr | nc localhost 2181 | grep -E "zk_server_state|zk_znode_count|zk_outstanding_requests"
# Step 3: Build the tree
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /
create /app "config-v1"
get /app
set /app "config-v2"
stat /app
ls -R /
quit
EOF
# stat: dataVersion increments; cZxid, mZxid
# Step 4: Make an ephemeral node
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -e /app/worker-1 "alive"
ls /app
quit
EOF
# reconnect: the session ended, so the ephemeral node is GONE
zkCli.sh -server localhost:2182 2>&1 <<'EOF'
ls /app
quit
EOF
# (asked of a different server: every one of them has the same tree)
# Step 5: Make sequential nodes
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -s /app/task- "t"
create -s /app/task- "t"
create -s /app/task- "t"
quit
EOF
# -> /app/task-0000000001, -0000000002, -0000000003
# [Corrected: this said -0000000000, -0000000001, -0000000002. The number is
# kept by the PARENT, /app, and counts every child it has had: worker-1 above
# took the first, though it is gone. Never assume a sequence starts at 0.]
# Step 6: Elect a leader
# 1. every candidate: create -e -s /election/n-
# 2. read the children; LOWEST sequence number is the leader
# 3. everyone else watches the node IMMEDIATELY BELOW their own
# -- not the leader. Watching the leader wakes every candidate on one
# failure: the HERD EFFECT.
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create /election ""
create -e -s /election/n- "nn1"
create -e -s /election/n- "nn2"
ls /election
get -w /election/n-0000000000
quit
EOF
# get -w sets a ONE-SHOT watch
# Step 7: Look for the systems that use it
zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /hbase
ls /hadoop-ha/mycluster
ls /rmstore
quit
EOF
# /hbase master, rs, meta-region-server
# /hadoop-ha/mycluster ActiveStandbyElectorLock <- NameNode HA
# /rmstore YARN ResourceManager HA
# Nothing uses this ensemble, so none of them exists here. Every HA story in
# the Hadoop ecosystem ends in znodes like these.
zkCli.sh -server localhost:2181 deleteall /app 2>/dev/null | tail -1
for id in 1 2 3; do zkServer.sh stop $PWD/zk$id.cfg 2>&1 | tail -1; done
# [Corrected: the zkCli.sh commands were indented under `zkCli.sh -server ...`,
# as typed at its prompt. Run as a script, the shell runs them itself -- `ls /`
# lists the machine's root directory, and `create` is "command not found".
# Each session is a here-document here, typed into zkCli.sh.]
The Python check, 16_zookeeper_model.py:
"""Experiment 16 -- demonstrate coordination with ZooKeeper.
`16_zookeeper.sh` carries the real `zkCli.sh` sessions, and runs on a real
three-server ensemble (the lab page shows it). What runs here is the
COORDINATION LOGIC: a
znode tree with ephemeral and sequential nodes, leader election by the
standard recipe, a distributed lock, and the quorum arithmetic that decides
whether an ensemble can make progress at all.
The point: ZooKeeper is not a database and not a queue. It is a small,
strongly-consistent tree whose only interesting properties are (a) ephemeral
nodes vanish when a session dies and (b) sequential nodes are numbered by a
single authority. Every recipe is built from exactly those two facts.
"""
class ZooKeeper:
def __init__(self):
self.tree = {"/": {"data": None, "ephemeral": False, "children": []}}
self.counter = {}
self.sessions = {}
self.watches = []
def create(self, path, data=None, ephemeral=False, sequential=False,
session=None):
parent = path.rsplit("/", 1)[0] or "/"
if parent not in self.tree:
raise KeyError(f"no node {parent} -- ZooKeeper creates no parents")
if sequential:
n = self.counter.get(path, 0)
self.counter[path] = n + 1
path = f"{path}{n:010d}"
if path in self.tree:
raise FileExistsError(f"{path} exists -- create is ATOMIC")
self.tree[path] = {"data": data, "ephemeral": ephemeral,
"children": [], "session": session}
self.tree[parent]["children"].append(path)
if ephemeral:
self.sessions.setdefault(session, []).append(path)
return path
def children(self, path):
return sorted(self.tree[path]["children"])
def expire(self, session):
"""A session dies -- every ephemeral node it owns disappears."""
gone = self.sessions.pop(session, [])
for p in gone:
parent = p.rsplit("/", 1)[0] or "/"
self.tree[parent]["children"].remove(p)
del self.tree[p]
self.watches.append(("deleted", p))
return gone
def quorum(n):
return n // 2 + 1
def main():
print(" Experiment 16 -- ZooKeeper coordination")
zk = ZooKeeper()
zk.create("/hadoop-ha")
zk.create("/hadoop-ha/mycluster")
# Step 1: Elect a leader
print("\n leader election, the standard recipe:")
print(" every candidate creates an EPHEMERAL SEQUENTIAL znode")
print(" the LOWEST sequence number is the leader")
print(" everyone else watches the node just below them")
nodes = {}
for host in ("nn1", "nn2", "nn3"):
p = zk.create("/hadoop-ha/mycluster/lock-", data=host,
ephemeral=True, sequential=True, session=host)
nodes[host] = p
print(f" {host} -> {p.rsplit('/', 1)[1]}")
order = zk.children("/hadoop-ha/mycluster")
leader = zk.tree[order[0]]["data"]
print(f"\n LEADER: {leader}")
assert leader == "nn1"
print("\n nn1's session expires (its JVM was killed):")
gone = zk.expire("nn1")
order = zk.children("/hadoop-ha/mycluster")
new_leader = zk.tree[order[0]]["data"]
print(f" {gone[0].rsplit('/', 1)[1]} vanished; new LEADER: {new_leader}")
assert new_leader == "nn2" and len(order) == 2
print(""" nobody ran a failover script. The ephemeral node was
deleted BY THE SERVER when the heartbeat stopped, the watch
fired, and nn2 saw itself at the head of the queue. That is
how HDFS NameNode HA actually chooses its active node --
which is the link back to experiment 5""")
print("\n why watch the node BELOW you, not the leader:")
print(f" {len(order) + 1} candidates all watching the leader means")
print(f" {len(order) + 1} clients woken by one failure -- the HERD EFFECT.")
print(""" Watching your immediate predecessor wakes exactly ONE
client per failure. The recipe is not arbitrary; it is a
thundering-herd fix, and examiners like that you know why""")
# Step 2: Take a distributed lock
print("\n a distributed lock is the SAME recipe:")
zk.create("/locks")
holders = []
for client in ("jobA", "jobB"):
p = zk.create("/locks/write-", data=client, ephemeral=True,
sequential=True, session=client)
holders.append((client, p))
first = zk.children("/locks")[0]
print(f" jobA and jobB both asked; {zk.tree[first]['data']} holds the lock")
assert zk.tree[first]["data"] == "jobA"
zk.expire("jobA")
nxt = zk.children("/locks")[0]
print(f" jobA CRASHES -- lock passes to {zk.tree[nxt]['data']} automatically")
assert zk.tree[nxt]["data"] == "jobB"
print(""" the lock is released by the SESSION DYING, not by the
client remembering to release it. A lock in a normal
database survives the crash of whoever held it and deadlocks
the system; an ephemeral znode cannot""")
# Step 3: Rely on an atomic create
print("\n create is ATOMIC, which is the other half of every recipe:")
zk.create("/config", data="v1")
try:
zk.create("/config", data="v2")
raise AssertionError("the second create must fail")
except FileExistsError as exc:
print(f" second create -> {type(exc).__name__}")
print(""" exactly one client wins a create, cluster-wide, with no
further negotiation. 'Whoever creates /master is the master'
is a complete election algorithm in one line, and it works
only because ZooKeeper linearises writes""")
# Step 4: Size the ensemble
print("\n ensemble sizing -- why every cluster has an ODD number:")
print(f" {'servers':>8}{'quorum':>8}{'can lose':>10} {'verdict'}")
for n in (1, 2, 3, 4, 5, 6, 7):
q = quorum(n)
tol = n - q
verdict = ("no fault tolerance" if tol == 0 else
"same tolerance as " + str(n - 1) if n % 2 == 0 else "good")
print(f" {n:>8}{q:>8}{tol:>10} {verdict}")
assert quorum(3) == 2 and quorum(4) == 3
assert (3 - quorum(3)) == (4 - quorum(4)) == 1
print(""" 3 servers tolerate 1 failure. FOUR SERVERS ALSO TOLERATE
ONE -- the extra machine buys nothing and adds a write to
every quorum. That is the whole reason ZooKeeper ensembles
are 3, 5 or 7, and it is a two-line exam answer""")
# Step 5: See what ZooKeeper is not
print("\n what ZooKeeper is NOT:")
print(f" {'misuse':<34}{'why it fails'}")
for m, w in (("a message queue", "no ordering guarantees across znodes"),
("a data store", "1 MB per znode, whole tree in RAM"),
("a cache", "every write is a quorum round trip"),
("a service registry for 10k nodes", "watch storms")):
print(f" {m:<34}{w}")
print(""" ZooKeeper stores COORDINATION STATE -- who is the leader,
who holds the lock, what is the config -- and it is small,
consistent and slow on purpose. Putting application data in
it is the mistake that gets clusters into trouble""")
# Step 6: Name who uses it
print("\n who uses it in this course:")
for who, why in (("HDFS NameNode HA", "elects the ACTIVE NameNode (exp 5)"),
("YARN ResourceManager HA", "elects the active RM (exp 6)"),
("HBase", "tracks the master and the RegionServers (exp 15)"),
("Kafka (pre-3.x)", "broker membership and controller")):
print(f" {who:<26}{why}")
print(""" every HA story in the Hadoop ecosystem ends at
ZooKeeper, which is why an experiment that looks like a
detour is actually the keystone""")
if __name__ == "__main__":
main()
On the cluster, 16_zookeeper.sh:
OUTPUT
$ for id in 1 2 3; do
mkdir -p zk$id && echo $id > zk$id/myid
cat > zk$id.cfg <<EOF
tickTime=2000
initLimit=10
syncLimit=5
dataDir=$PWD/zk$id
clientPort=218$id
server.1=localhost:2888:3888
server.2=localhost:2889:3889
server.3=localhost:2890:3890
4lw.commands.whitelist=srvr,mntr
EOF
done
$ for id in 1 2 3; do ZOO_LOG_DIR=zk$id zkServer.sh start $PWD/zk$id.cfg 2>&1 | tail -1; done
Starting zookeeper ... STARTED
Starting zookeeper ... STARTED
Starting zookeeper ... STARTED
$ sleep 10
$ for id in 1 2 3; do zkServer.sh status $PWD/zk$id.cfg 2>&1 | grep Mode; done
Mode: follower
Mode: leader
Mode: follower
$ echo srvr | nc localhost 2181 | grep -E "Mode|Zxid|Node count"
Zxid: 0x0
Mode: follower
Node count: 5
$ echo mntr | nc localhost 2181 | grep -E "zk_server_state|zk_znode_count|zk_outstanding_requests"
zk_server_state follower
zk_outstanding_requests 0
zk_znode_count 5
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /
create /app "config-v1"
get /app
set /app "config-v2"
stat /app
ls -R /
quit
EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] ls /
[zookeeper]
[zk: localhost:2181(CONNECTED) 1] create /app "config-v1"
Created /app
[zk: localhost:2181(CONNECTED) 2] get /app
config-v1
[zk: localhost:2181(CONNECTED) 3] set /app "config-v2"
[zk: localhost:2181(CONNECTED) 4] stat /app
cZxid = 0x100000002
ctime = Sun Oct 04 23:31:34 UTC 2026
mZxid = 0x100000003
mtime = Sun Oct 04 23:31:34 UTC 2026
pZxid = 0x100000002
cversion = 0
dataVersion = 1
aclVersion = 0
ephemeralOwner = 0x0
dataLength = 9
numChildren = 0
[zk: localhost:2181(CONNECTED) 5] ls -R /
/
/app
/zookeeper
/zookeeper/config
/zookeeper/quota
[zk: localhost:2181(CONNECTED) 6] quit
WATCHER::
WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -e /app/worker-1 "alive"
ls /app
quit
EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create -e /app/worker-1 "alive"
Created /app/worker-1
[zk: localhost:2181(CONNECTED) 1] ls /app
[worker-1]
[zk: localhost:2181(CONNECTED) 2] quit
WATCHER::
WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2182 2>&1 <<'EOF'
ls /app
quit
EOF
Connecting to localhost:2182
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2182(CONNECTED) 0] ls /app
[]
[zk: localhost:2182(CONNECTED) 1] quit
WATCHER::
WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create -s /app/task- "t"
create -s /app/task- "t"
create -s /app/task- "t"
quit
EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create -s /app/task- "t"
Created /app/task-0000000001
[zk: localhost:2181(CONNECTED) 1] create -s /app/task- "t"
Created /app/task-0000000002
[zk: localhost:2181(CONNECTED) 2] create -s /app/task- "t"
Created /app/task-0000000003
[zk: localhost:2181(CONNECTED) 3] quit
WATCHER::
WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
create /election ""
create -e -s /election/n- "nn1"
create -e -s /election/n- "nn2"
ls /election
get -w /election/n-0000000000
quit
EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] create /election ""
Created /election
[zk: localhost:2181(CONNECTED) 1] create -e -s /election/n- "nn1"
Created /election/n-0000000000
[zk: localhost:2181(CONNECTED) 2] create -e -s /election/n- "nn2"
Created /election/n-0000000001
[zk: localhost:2181(CONNECTED) 3] ls /election
[n-0000000000, n-0000000001]
[zk: localhost:2181(CONNECTED) 4] get -w /election/n-0000000000
nn1
[zk: localhost:2181(CONNECTED) 5] quit
WATCHER::
WatchedEvent state:SyncConnected type:NodeDeleted path:/election/n-0000000000
WATCHER::
WatchedEvent state:Closed type:None path:null
$ zkCli.sh -server localhost:2181 2>&1 <<'EOF'
ls /hbase
ls /hadoop-ha/mycluster
ls /rmstore
quit
EOF
Connecting to localhost:2181
Welcome to ZooKeeper!
JLine support is enabled
WATCHER::
WatchedEvent state:SyncConnected type:None path:null
[zk: localhost:2181(CONNECTED) 0] ls /hbase
Node does not exist: /hbase
[zk: localhost:2181(CONNECTED) 1] ls /hadoop-ha/mycluster
Node does not exist: /hadoop-ha/mycluster
[zk: localhost:2181(CONNECTED) 2] ls /rmstore
Node does not exist: /rmstore
[zk: localhost:2181(CONNECTED) 3] quit
WATCHER::
WatchedEvent state:Closed type:None path:null
Exiting JVM with code 1
$ zkCli.sh -server localhost:2181 deleteall /app 2>/dev/null | tail -1
WatchedEvent state:SyncConnected type:None path:null
$ for id in 1 2 3; do zkServer.sh stop $PWD/zk$id.cfg 2>&1 | tail -1; done
Stopping zookeeper ... STOPPED
Stopping zookeeper ... STOPPED
Stopping zookeeper ... STOPPED
The Python check, 16_zookeeper_model.py:
OUTPUT
Experiment 16 -- ZooKeeper coordination
leader election, the standard recipe:
every candidate creates an EPHEMERAL SEQUENTIAL znode
the LOWEST sequence number is the leader
everyone else watches the node just below them
nn1 -> lock-0000000000
nn2 -> lock-0000000001
nn3 -> lock-0000000002
LEADER: nn1
nn1's session expires (its JVM was killed):
lock-0000000000 vanished; new LEADER: nn2
nobody ran a failover script. The ephemeral node was
deleted BY THE SERVER when the heartbeat stopped, the watch
fired, and nn2 saw itself at the head of the queue. That is
how HDFS NameNode HA actually chooses its active node --
which is the link back to experiment 5
why watch the node BELOW you, not the leader:
3 candidates all watching the leader means
3 clients woken by one failure -- the HERD EFFECT.
Watching your immediate predecessor wakes exactly ONE
client per failure. The recipe is not arbitrary; it is a
thundering-herd fix, and examiners like that you know why
a distributed lock is the SAME recipe:
jobA and jobB both asked; jobA holds the lock
jobA CRASHES -- lock passes to jobB automatically
the lock is released by the SESSION DYING, not by the
client remembering to release it. A lock in a normal
database survives the crash of whoever held it and deadlocks
the system; an ephemeral znode cannot
create is ATOMIC, which is the other half of every recipe:
second create -> FileExistsError
exactly one client wins a create, cluster-wide, with no
further negotiation. 'Whoever creates /master is the master'
is a complete election algorithm in one line, and it works
only because ZooKeeper linearises writes
ensemble sizing -- why every cluster has an ODD number:
servers quorum can lose verdict
1 1 0 no fault tolerance
2 2 0 no fault tolerance
3 2 1 good
4 3 1 same tolerance as 3
5 3 2 good
6 4 2 same tolerance as 5
7 4 3 good
3 servers tolerate 1 failure. FOUR SERVERS ALSO TOLERATE
ONE -- the extra machine buys nothing and adds a write to
every quorum. That is the whole reason ZooKeeper ensembles
are 3, 5 or 7, and it is a two-line exam answer
what ZooKeeper is NOT:
misuse why it fails
a message queue no ordering guarantees across znodes
a data store 1 MB per znode, whole tree in RAM
a cache every write is a quorum round trip
a service registry for 10k nodes watch storms
ZooKeeper stores COORDINATION STATE -- who is the leader,
who holds the lock, what is the config -- and it is small,
consistent and slow on purpose. Putting application data in
it is the mistake that gets clusters into trouble
who uses it in this course:
HDFS NameNode HA elects the ACTIVE NameNode (exp 5)
YARN ResourceManager HA elects the active RM (exp 6)
HBase tracks the master and the RegionServers (exp 15)
Kafka (pre-3.x) broker membership and controller
every HA story in the Hadoop ecosystem ends at
ZooKeeper, which is why an experiment that looks like a
detour is actually the keystone
The script starts and stops its own ensemble, three servers on one machine. Which of them wins
the election differs from run to run; exactly one says Mode: leader.
WATCH YOUR PREDECESSOR, NOT THE LEADER
Watching the leader wakes every candidate on one failure — the herd effect. Watching your immediate predecessor wakes exactly one.
The lock is the same recipe.
jobA holds the lock
jobA CRASHES -- lock passes to jobB automatically
Released by the session dying, not by the client remembering. A database lock survives its holder's crash and deadlocks the system; an ephemeral znode cannot.
ENSEMBLE SIZING
| Servers | Quorum | Can lose | Verdict |
|---|---|---|---|
| 3 | 2 | 1 | good |
| 4 | 3 | 1 | same as 3 |
| 5 | 3 | 2 | good |
| 6 | 4 | 2 | same as 5 |
| 7 | 4 | 3 | good |
Four servers tolerate one failure — exactly what three tolerate. The fourth machine buys nothing and adds a write to every quorum. 3, 5 or 7.
RESULT
Three servers elected one leader; every server held the same tree; an ephemeral node vanished when its session ended; the sequence numbers started at 1, not 0, because the parent counts every child it has had. In the model, four servers tolerate one failure, as three do.
Process HBase datasets using Spark integration with Hadoop.
Read an HBase table into Spark as an RDD, a partition per region, and query it with Spark SQL; then, in PySpark, count words with RDDs, compare reduceByKey with groupByKey, query the star schema and cache.
On the cluster, 17_spark_hbase.scala:
The Python check, 17_spark.py:
TWO REAL ENGINES
real SparkSession: version 4.2.0, master local[2]
PySpark installs from PyPI and Java 21 is present, so a genuine session
starts, real RDDs are built, and a real shuffle happens inside
reduceByKey. The Scala half runs in spark-shell against HBase, with HBase's
own jars on the classpath and its TableInputFormat — the connector Hadoop ships —
so no separate connector is needed.
On the cluster, 17_spark_hbase.scala:
// Experiment 17 -- process HBase datasets using Spark integration with Hadoop
//
// Run it: the spark-shell command below, with HBase running. It was run on a Hadoop 3.3.6 cluster where these labs
// are checked (tools/data-science/hadoop_lab.py), and the lab page shows what it printed.
// [Changed: this said the file had never been run, as the Hadoop stack could not be
// installed there. It installs from archive.apache.org: tools/data-science/setup_hadoop.sh.]
//
// The runnable half is 17_spark.py, which runs REAL PySpark -- only the HBase connector is missing
//
// run with:
// spark-shell --master 'local[*]' \
// --jars $(ls $HBASE_HOME/lib/*.jar | tr '\n' ',') \
// -i 17_spark_hbase.scala
// [Corrected: the jars were in /usr/lib/hbase/lib, where some distributions put
// HBase; $HBASE_HOME is wherever yours is. `hbase mapredcp` lists the smaller
// set a job needs, and works the same. And --master yarn needs YARN containers
// running Java 17 for Spark 4; this cluster's run Java 8, so this runs
// local[*] -- the same code, on one machine.]
import org.apache.hadoop.hbase.{HBaseConfiguration, CellUtil}
import org.apache.hadoop.hbase.client.Result
import org.apache.hadoop.hbase.io.ImmutableBytesWritable
import org.apache.hadoop.hbase.mapreduce.TableInputFormat
import org.apache.hadoop.hbase.util.Bytes
// Step 1: Configure the HBase connection and scan
val conf = HBaseConfiguration.create()
conf.set("hbase.zookeeper.quorum", "localhost") // ZooKeeper, again
// [Corrected: the quorum was zk1,zk2,zk3, experiment 16's three hosts. HBase
// here runs its own ZooKeeper, on localhost.]
conf.set(TableInputFormat.INPUT_TABLE, "sales")
// push the scan down: read one region, not the table
conf.set(TableInputFormat.SCAN_ROW_START, "South")
conf.set(TableInputFormat.SCAN_ROW_STOP, "South~")
conf.set(TableInputFormat.SCAN_COLUMNS, "sales:revenue sales:qty")
// Step 2: Read the table as an RDD, a partition per region
val hBaseRDD = sc.newAPIHadoopRDD(
conf,
classOf[TableInputFormat],
classOf[ImmutableBytesWritable],
classOf[Result])
// ONE SPARK PARTITION PER HBASE REGION. That is the whole integration:
// Spark reads regions in parallel, locally, without going through the
// RegionServer's RPC path for bulk scans.
println(s"partitions = ${hBaseRDD.getNumPartitions}")
// [Noted, from running it: straight after the puts this said partitions = 0,
// and every query below came back empty, with no error. The rows were still
// in the memstore, in no store file, so the region's size was 0 -- and
// TableInputFormat made no split for it. `flush 'sales'` in the HBase shell
// first; on a busy table the memstore flushes by itself.]
// Step 3: Map each row to a case class
case class Sale(rowKey: String, region: String, qty: Int, revenue: Double)
val sales = hBaseRDD.map { case (_, result) =>
val key = Bytes.toString(result.getRow)
val qty = Option(result.getValue(Bytes.toBytes("sales"), Bytes.toBytes("qty")))
.map(b => Bytes.toString(b).toInt).getOrElse(0)
val rev = Option(result.getValue(Bytes.toBytes("sales"), Bytes.toBytes("revenue")))
.map(b => Bytes.toString(b).toDouble).getOrElse(0.0)
Sale(key, key.split("#")(0), qty, rev)
}
// Step 4: Query it with Spark SQL
import spark.implicits._
val df = sales.toDF()
df.createOrReplaceTempView("sales")
spark.sql("""
SELECT region, SUM(revenue) AS revenue, SUM(qty) AS units
FROM sales GROUP BY region ORDER BY revenue DESC
""").show()
// expected: South 10360 -- the same number Course 11's DAX, experiment 10's
// SQL and experiment 17's PySpark all produce.
// [Corrected: this expected North 2520 as well. The scan above reads only the
// rows from 'South' to 'South~' -- that is the point of pushing it down -- so
// North is never read, and the query has one row.]
// --- writing BACK to HBase, in bulk ---------------------------------------
// Never use put() per row from a Spark job: that is one RPC per record and it
// will overwhelm the RegionServers. Write HFiles and load them:
//
// df.rdd.map(toKeyValue).sortByKey()
// .saveAsNewAPIHadoopFile(path, classOf[ImmutableBytesWritable],
// classOf[KeyValue], classOf[HFileOutputFormat2], conf)
// LoadIncrementalHFiles.doBulkLoad(new Path(path), admin, table, locator)
//
// Bulk load bypasses the write path entirely -- no WAL, no memstore, no
// flush -- and is one to two orders of magnitude faster than put().
// --- when NOT to do this ---------------------------------------------------
// A full-table Spark scan of HBase is SLOWER than the same data in Parquet,
// because HBase stores every cell with its row key, family, qualifier and
// timestamp. HBase is for random reads and writes; Parquet is for scans.
// If every job you run is a full scan, the data is in the wrong store.
The Python check, 17_spark.py:
"""Experiment 17 -- process HBase datasets using Spark integration with Hadoop.
THIS EXPERIMENT RUNS REAL SPARK. PySpark 4.2 installs from PyPI and Java 21 is
present, so a genuine SparkSession starts, real RDDs are built, and a real
shuffle happens inside reduceByKey. Nothing here is a simulation.
The dataset here comes from the experiment 15 model rather than an HBase table;
`17_spark_hbase.scala` reads a real one, through TableInputFormat, in
spark-shell (the lab page shows it).
Run with: /tmp/sparkenv/bin/python 17_spark.py
or let tools/run_bigdata_labs.py find the environment for you.
"""
import os
import sys
import fixtures as f
def spark_available():
try:
import pyspark # noqa: F401
return True
except ImportError:
return False
def main():
print(" Experiment 17 -- Spark on the Hadoop stack")
if not spark_available():
print("""
PySpark is not importable from this interpreter.
Run tools/setup_spark.sh, then use /tmp/sparkenv/bin/python.
SKIPPED -- and this line is what a skipped experiment looks like.""")
return False
os.environ.setdefault("PYSPARK_PYTHON", sys.executable)
from pyspark.sql import SparkSession
from pyspark.sql import functions as F
# Step 1: Start a SparkSession
spark = (SparkSession.builder
.appName("course-12b-exp-17")
.master("local[2]")
.config("spark.ui.enabled", "false")
.config("spark.sql.shuffle.partitions", "4")
.getOrCreate())
spark.sparkContext.setLogLevel("ERROR")
print(f"\n real SparkSession: version {spark.version}, "
f"master {spark.sparkContext.master}")
# Step 2: Count words with RDDs
rdd = spark.sparkContext.parallelize(list(f.DOCS.values()), 3)
counts = (rdd.flatMap(lambda line: line.split())
.map(lambda w: (w, 1))
.reduceByKey(lambda a, b: a + b))
got = dict(counts.collect())
print(f"\n RDD word count: {len(got)} distinct words, "
f"{sum(got.values())} total")
assert sum(got.values()) == 48
assert got["the"] == 5 and got["big"] == 4 and got["dog"] == 4
print(f" partitions: input {rdd.getNumPartitions()}, "
f"after reduceByKey {counts.getNumPartitions()}")
print(""" IDENTICAL to experiment 7's MapReduce answer, on a real
distributed engine. reduceByKey is map -> COMBINE -> shuffle
-> reduce; Spark applies the combiner automatically, which
MapReduce makes you ask for""")
# Step 3: Compare reduceByKey with groupByKey
print("\n reduceByKey against groupByKey -- the same answer, not the same job:")
grouped = (rdd.flatMap(lambda line: line.split())
.map(lambda w: (w, 1))
.groupByKey()
.mapValues(len))
assert dict(grouped.collect()) == got
# measure the map-side combine for THIS partitioning rather than quoting
# experiment 7's number, which was per-document and not per-partition
per_part = (rdd.flatMap(lambda line: line.split())
.map(lambda w: (w, 1))
.mapPartitions(lambda it: [len({k for k, _ in it})])
.collect())
combined = sum(per_part)
print(f" groupByKey : shuffles all 48 pairs, then counts")
print(f" reduceByKey : combines to {combined} map-side "
f"({per_part} per partition), then shuffles")
assert combined < 48
print(""" same output, and groupByKey moves every record across
the network while reduceByKey moves one per key per
partition. On a real corpus groupByKey is how you produce an
OutOfMemoryError on a single hot key. This is the most
examined Spark question there is""")
# Step 4: See lazy evaluation and the DAG
lineage = (rdd.flatMap(lambda l: l.split())
.filter(lambda w: len(w) > 3)
.map(lambda w: (w[0], 1)))
print(f"\n lazy evaluation: three transformations queued, nothing ran")
print(f" the DAG has {len(lineage.toDebugString().decode().splitlines())} "
f"stages of lineage recorded")
result = lineage.reduceByKey(lambda a, b: a + b).collect()
print(f" .collect() is the ACTION -- it returned {len(result)} keys")
assert len(result) > 0
print(""" transformations build a DAG; only an ACTION submits it.
That is why a typo in a map() surfaces at collect() and not
where you wrote it -- and why Spark can fuse the whole chain
into one pass over the data""")
# Step 5: Query the star schema with DataFrames
sdf = spark.createDataFrame(f.SALES_DF)
agg = (sdf.groupBy("region")
.agg(F.sum("revenue").alias("revenue"),
F.sum("profit").alias("profit"))
.orderBy(F.desc("revenue")))
rows = {r["region"]: r["revenue"] for r in agg.collect()}
print(f"\n DataFrame aggregate over the SAME nine rows:")
for r in agg.collect():
print(f" {r['region']:<8}{r['revenue']:>10,.0f}{r['profit']:>9,.0f}")
assert rows["South"] == 10360.0 and rows["North"] == 2520.0
assert sum(rows.values()) == f.total_revenue()
print(""" 10,360 and 2,520 again -- the third engine to produce
them, after Course 11's DAX and experiment 10's SQL. Spark,
DuckDB and Power BI agree, which is what reusing one dataset
across three courses was for""")
# Step 6: Analyse the logs
logs = spark.sparkContext.parallelize(f.access_logs(40), 2)
by_status = (logs.map(lambda ln: (ln.rsplit(" ", 2)[-2], 1))
.reduceByKey(lambda a, b: a + b)
.collectAsMap())
print(f"\n the ingested access logs, aggregated in Spark:")
for code in sorted(by_status):
print(f" HTTP {code}: {by_status[code]}")
assert by_status["200"] == 24 and by_status["404"] == 8
assert sum(by_status.values()) == 40
print(""" the same 24 / 8 / 8 the Flume agent produced in
experiment 12. Ingest with Flume, analyse with Spark, on
bytes that were never transformed in between -- that is the
end-to-end story the syllabus asks for""")
# Step 7: Compare Spark with MapReduce
print("\n Spark against MapReduce, on the parts that decided it:")
print(f" {'':<22}{'MapReduce':<26}{'Spark'}")
for label, mr, sp in (
("between stages", "writes to HDFS", "keeps in MEMORY"),
("iterative jobs", "re-reads every pass", "cache() once"),
("API", "map and reduce only", "~80 operators"),
("interactive", "no", "yes -- the shell"),
("fault tolerance", "re-run the task", "recompute from LINEAGE"),
("streaming", "no", "structured streaming"),
("runs on YARN", "yes", "yes -- same cluster")):
print(f" {label:<22}{mr:<26}{sp}")
print(""" the decisive row is the first. A ten-iteration machine
learning job writes to HDFS nine times under MapReduce and
zero times under Spark, which is where the '100x faster'
headline comes from -- it is a claim about ITERATIVE jobs,
and quoting it for a single-pass job is wrong""")
# Step 8: Cache, and measure it
base = spark.sparkContext.parallelize(range(200_000), 4).map(lambda x: x * 2)
base.cache()
first = base.sum()
second = base.sum()
assert first == second == sum(x * 2 for x in range(200_000))
print(f"\n cache(): two actions over the same RDD, sum = {first:,}")
print(f" storage level after cache(): {base.getStorageLevel()}")
print(""" WITHOUT cache() the second sum recomputes the map from
the source. With it, only the first action pays. Caching is
the single highest-value Spark optimisation and the one
students forget, because nothing FAILS without it -- the job
is merely twice as slow""")
base.unpersist()
spark.stop()
print("\n SparkSession stopped cleanly.")
return True
if __name__ == "__main__":
main()
On the cluster, 17_spark_hbase.scala:
OUTPUT
$ start-hbase.sh > /dev/null 2>&1
$ sleep 20
$ hbase shell -n < load_sales.hbase 2>&1 | grep -E 'row\(s\)$|^=> [0-9]+$'
9 row(s)
=> 9
$ JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64 $SPARK_HOME/bin/spark-shell --master 'local[*]' --conf spark.ui.enabled=false --jars $(ls $HBASE_HOME/lib/*.jar | tr '\n' ',') -i 17_spark_hbase.scala < /dev/null 2>spark.log
Welcome to
____ __
/ __/__ ___ _____/ /__
_\ \/ _ \/ _ `/ __/ '_/
/___/ .__/\_,_/_/ /_/\_\ version 4.2.0
/_/
Using Scala version 2.13.18 (OpenJDK 64-Bit Server VM, Java 21.0.10)
Type in expressions to have them evaluated.
Type :help for more information.
Spark context available as 'sc' (master = local[*], app id = local-1791156759020).
Spark session available as 'spark'.
partitions = 1
+------+-------+-----+
|region|revenue|units|
+------+-------+-----+
| South|10360.0| 48|
+------+-------+-----+
scala> :quit
$ awk '/ERROR|Exception/' spark.log | sed -E 's/^[0-9/]+ [0-9:]+ //' | sort -u | head -5
$ stop-hbase.sh 2>&1 | tail -1
stopping hbase.............
The Python check, 17_spark.py:
OUTPUT
Experiment 17 -- Spark on the Hadoop stack
real SparkSession: version 4.2.0, master local[2]
RDD word count: 26 distinct words, 48 total
partitions: input 3, after reduceByKey 3
IDENTICAL to experiment 7's MapReduce answer, on a real
distributed engine. reduceByKey is map -> COMBINE -> shuffle
-> reduce; Spark applies the combiner automatically, which
MapReduce makes you ask for
reduceByKey against groupByKey -- the same answer, not the same job:
groupByKey : shuffles all 48 pairs, then counts
reduceByKey : combines to 35 map-side ([11, 14, 10] per partition), then shuffles
same output, and groupByKey moves every record across
the network while reduceByKey moves one per key per
partition. On a real corpus groupByKey is how you produce an
OutOfMemoryError on a single hot key. This is the most
examined Spark question there is
lazy evaluation: three transformations queued, nothing ran
the DAG has 2 stages of lineage recorded
.collect() is the ACTION -- it returned 11 keys
transformations build a DAG; only an ACTION submits it.
That is why a typo in a map() surfaces at collect() and not
where you wrote it -- and why Spark can fuse the whole chain
into one pass over the data
DataFrame aggregate over the SAME nine rows:
South 10,360 2,760
North 2,520 765
10,360 and 2,520 again -- the third engine to produce
them, after Course 11's DAX and experiment 10's SQL. Spark,
DuckDB and Power BI agree, which is what reusing one dataset
across three courses was for
the ingested access logs, aggregated in Spark:
HTTP 200: 24
HTTP 404: 8
HTTP 500: 8
the same 24 / 8 / 8 the Flume agent produced in
experiment 12. Ingest with Flume, analyse with Spark, on
bytes that were never transformed in between -- that is the
end-to-end story the syllabus asks for
Spark against MapReduce, on the parts that decided it:
MapReduce Spark
between stages writes to HDFS keeps in MEMORY
iterative jobs re-reads every pass cache() once
API map and reduce only ~80 operators
interactive no yes -- the shell
fault tolerance re-run the task recompute from LINEAGE
streaming no structured streaming
runs on YARN yes yes -- same cluster
the decisive row is the first. A ten-iteration machine
learning job writes to HDFS nine times under MapReduce and
zero times under Spark, which is where the '100x faster'
headline comes from -- it is a claim about ITERATIVE jobs,
and quoting it for a single-pass job is wrong
cache(): two actions over the same RDD, sum = 39,999,800,000
storage level after cache(): Memory Serialized 1x Replicated
WITHOUT cache() the second sum recomputes the map from
the source. With it, only the first action pays. Caching is
the single highest-value Spark optimisation and the one
students forget, because nothing FAILS without it -- the job
is merely twice as slow
SparkSession stopped cleanly.
_drive_17_spark_hbase.py starts HBase as for experiment 15, puts the nine sales rows into
sales with the row key region#store#date#product, and flushes the table — a scan through
TableInputFormat reads the store files, and the rows still in memory were not seen until the
flush; the file says so. spark-shell runs on Java 21 and HBase on Java 8.
17_spark.py's output is what it printed to stdout. Spark's JVM logs to stderr — a timestamped
line per event and a progress bar — and that is left out. Among it is one warning worth knowing:
PySpark 4.2 "does not yet fully support pandas >= 3.0.0". The program uses pandas only to load
the shared fixtures, and every figure it prints is asserted.
The RDD word count. 26 distinct words, 48 total — identical to experiment 7's MapReduce answer, on a real distributed engine.
REDUCEBYKEY AGAINST GROUPBYKEY
| What crosses the network | |
|---|---|
groupByKey |
all 48 pairs, then counts |
reduceByKey |
combines to 35 map-side — [11, 14, 10] per partition |
WHY IT MATTERS
Note 35, not experiment 7's 39. Spark's three partitions each hold two documents, so more merging happens per task. The combiner's saving depends on the split, exactly as Unit 3 said — and this is the same measurement made two ways.
Lazy evaluation. Three transformations queued; .collect() is the action that submits
the DAG. That is why a typo in a map() surfaces at collect().
THE CROSS-COURSE CHECK, FOURTH ENGINE
| Region | Revenue | Profit |
|---|---|---|
| South | 10,360 | 2,760 |
| North | 2,520 | 765 |
Business Intelligence Tools' DAX, DuckDB, Hive and Spark all produce ₹10,360.
The logs from experiment 12. HTTP 200: 24, 404: 8, 500: 8 — the same 24 / 8 / 8 the Flume
agent produced. Ingest with Flume, analyse with Spark, on bytes never transformed in between.
cache().
two actions over the same RDD, sum = 39,999,800,000
storage level: Memory Serialized 1x Replicated
Without cache() the second action recomputes the whole lineage. Caching
is the highest-value Spark optimisation and the one students forget —
because nothing fails without it; the job is merely twice as slow.
THE SPARK-OVER-HBASE CAVEAT
One Spark partition per HBase region is the whole integration — the nine rows are one
region, so one partition. But never put() per row from a Spark job — write HFiles and
bulk-load them.
And: a full-table Spark scan of HBase is slower than the same data in Parquet, because HBase stores every cell with its row key, family, qualifier and timestamp. If every job is a full scan, the data is in the wrong store.
RESULT
Spark read the nine rows from HBase in one partition — one region — and Spark SQL gave South ₹10,360 from 48 units, as DAX, Hive and DuckDB did. The scan saw nothing until the table was flushed. In PySpark the word count is 26 words and 48 in all, and reduceByKey shuffles 35 records against groupByKey's 48.
| Script | Experiments | Real tool? |
|---|---|---|
04_blocks_replication.py |
4 | arithmetic |
05_fault_tolerance.py |
5 | model |
06_yarn_scheduling.py |
6 | model |
07_wordcount.py |
7 | explicit MapReduce engine |
08_inverted_index.py |
8 | same engine |
09_pig_equivalent.py |
9 | dataflow, step by step |
10_hive_duckdb.py |
10 | real SQL, DuckDB |
11_sqoop_equivalent.py |
11 | real SQLite + real Parquet |
12_flume_equivalent.py |
12 | agent semantics |
13_avro_parquet.py |
13 | real Avro + real Parquet |
14_pipeline.py |
14 | real end-to-end |
15_hbase_model.py |
15 | model |
16_zookeeper_model.py |
16 | model |
17_spark.py |
17 | REAL APACHE SPARK |
Plus the 15 tool files, run on the cluster, each checked for the answers it must print — Pig's three categories, Hive's South ₹10,360, Sqoop's 90 imported and 3 exported, Flume's 8 routed events — and none may still say NOT EXECUTED.
Experiment 17's PySpark half skips loudly if the PySpark environment is absent, and the tool files are only audited if the Hadoop stack is — the same graceful-skip pattern Web Technologies uses for jsdom. A skip is not a pass, and the runner says so.
Two hours on a cluster, one experiment number, then a viva.
What costs marks:
Iterable twice--split-by on a skewed or text columnexec tail -F as a Flume sourceWhat earns them:
The small-files factor: 222,222. One number that justifies HDFS's whole design.
"Any two failures survive; only some threes are fatal." 6 of 20, not "it breaks at three".
"104 container-seconds either way." Scheduling moves latency, it does not create throughput.
"The combiner saved 18.75% here because the splits are tiny." Naming why a result is unimpressive is stronger than quoting an impressive one.
Reporting the empty bucket. Hashing three keys into three buckets left one empty, and saying so beats pretending otherwise.
The three compression ratios: 0.2×, 303×, 12× — and saying which one to quote and why.
"Aggregate to a common grain, then join." Nine words that prevent the fan trap.
"You cannot have even write distribution and range scans." The one sentence of HBase row-key design.
"3, 5 or 7 — four tolerates what three does."
The same experiments, one page each, so a program can be reached by what it does rather than by its number.