Part of the data-platform path: Big Data Technologies, Cloud Computing for Data Science, Time Series Analysis and Forecasting and Data Engineering and MLOps.
"Big data" is not a size. It is the point at which the data does not fit on one machine, and everything you know stops working.
That threshold moves. In 2006 it was a few hundred gigabytes; a laptop now holds two terabytes and a single cloud VM can be rented with 24 TB of RAM. So the useful definition is not "more than N bytes" — it is:
NOTE
Data is big when the cost of moving it exceeds the cost of computing on it.
Every design decision in this course follows from that one sentence.
| The old assumption | What replaces it | Where you see it |
|---|---|---|
| Move data to the code | Move the code to the data | HDFS data locality, Unit 2 |
| Disks are reliable | Disks fail constantly; replicate | Replication factor 3, Unit 2 |
| One machine, scale UP | Many machines, scale OUT | YARN, Unit 2 |
| Schema before data | Schema when you read it | Hive external tables, Unit 3 |
| Update rows in place | Write once, append only | HDFS, Unit 2 |
| A transaction is atomic | Eventual consistency is enough | HBase, Unit 5 |
This course teaches a stack that peaked around 2015. MapReduce has been superseded by Spark; Sqoop and Flume were retired to the Apache Attic in 2021; Hadoop-on-premises is losing to object storage plus a query engine.
That is not a reason to skip it, and here is why:
The problems have not changed. Partitioning, shuffling, skew, data locality, the small-files problem, schema evolution and the batch/stream split are exactly the same in Spark, BigQuery, Snowflake and Databricks. Hadoop is where they are visible, because Hadoop makes you do them by hand.
The vocabulary is universal. "Partition", "shuffle", "predicate pushdown", "columnar format" mean the same thing everywhere.
Parquet, Avro and Spark are current. Three of the things this course teaches are what the industry actually uses today, and this repository runs all three for real.
So learn the concepts as permanent and the tools as historical, and say so in the exam — an answer that places Hadoop in time reads as understanding rather than memorisation.
Everything runs.
| The tools, on a Hadoop 3.3.6 cluster | The checks, in Python |
|---|---|
HDFS and YARN — hdfs dfs, fsck, dfsadmin, the sample jobs, the queues |
A MapReduce engine written out in full, with a visible shuffle |
Java MapReduce — WordCount.java and InvertedIndex.java, compiled and run |
DuckDB for Hive-style SQL |
| Pig and Hive | SQLite as the RDBMS for the Sqoop import |
| Sqoop, from MariaDB, and a Flume agent | Avro via fastavro and Parquet via pyarrow — real files |
| HBase, a three-server ZooKeeper ensemble, and Spark reading HBase | Apache Spark — a genuine SparkSession, real RDDs, a real shuffle |
tools/data-science/run_bigdata_labs.py runs both halves and checks the answers each gives;
lab.md shows what every file printed.
Updated October 2026: Hadoop and its ecosystem could not be installed where these notes are
checked, and every tool file said NOT EXECUTED. They now install from archive.apache.org
(tools/data-science/setup_hadoop.sh). Running them found what reading them had not, and
twelve of the fifteen files now carry corrections, each noted in the file and in
lab.md.
Introduce students to the concepts, characteristics, and challenges of Big Data.
Familiarize students with the Hadoop ecosystem and its core components (HDFS, YARN, MapReduce).
Develop practical knowledge of distributed storage and parallel processing in Hadoop.
Provide hands-on exposure to data ingestion tools (Sqoop, Flume) and serialization techniques.
Enable students to explore NoSQL databases (HBase), coordination services (ZooKeeper), and HadoopSpark integration for large-scale data analysis.
WATCH OUT
Objective 5 and Outcome 5 both read HadoopSpark, with the space lost. They mean Hadoop and Spark — two systems, and the distinction is the whole point of Unit 5. See review finding D24.
| Unit | Question it answers |
|---|---|
| 1 | What is big data, and what is in the Hadoop box? |
| 2 | Where does the data live, and who decides what runs? |
| 3 | How do you compute over it? |
| 4 | How does it get in, and in what format? |
| 5 | What if you need random access, coordination or speed? |
Read them in order. Unit 2 is the load-bearing one: HDFS's design decisions explain almost everything else in the course, and a student who understands blocks, replication and the NameNode's memory can derive most of Units 3–5.
labs/course-12b-bigdata/ — the code, and the runner that asserts every figure
these notes quote
data/course-12b-bigdata/ — practice datasets, CSV: web-logs.csv, wordcount-corpus.csv.
Every one was generated from a known truth, so you can score your answer
rather than just produce one; data/README.md lists what each was built
from, data/PRACTICE-QUESTIONS.md sets questions on each with a computed
answer key, and tools/data-science/check_datasets.py proves every one of those
answers against the file.
Also sales-transactions.csv in data/shared/, which several courses
analyse so their answers can be compared.
This course does not stand alone, and the labs make the links checkable.
| From | To | What is shared |
|---|---|---|
| Business Intelligence Tools (BI) | experiments 10, 14, 17 | The same nine-row star schema, imported not copied. South = ₹10,360 is produced by DAX, DuckDB and Spark, and the suite fails if they disagree. |
| Database Management Systems (DBMS) | Units 3 and 5 | Hive is SQL that compiles to a job; HBase is what you get when you drop joins, indexes and transactions. |
| Document Oriented Database (MongoDB) | Unit 5 | Two NoSQL stores, compared directly. The sharpest difference: MongoDB has secondary indexes; HBase does not. |
| Python for Data Analysis and Visualization | throughout | pandas is the single-machine version of every operation here. |
| Machine Learning (ML) | Unit 3 | Spark's MLlib is the same algorithms, distributed — and the reason Spark beat MapReduce is iterative ML jobs. |
White, Hadoop: The Definitive Guide, 4th edition, O'Reilly — the reference for Units 1–3. The syllabus gives "4th" and then stops, without a publisher.
Damji, Wenig, Das & Lee, Learning Spark, 2nd edition, O'Reilly, 2020 — Unit 5, and free from Databricks.
References: BIG DATA, Black Book, DreamTech Press, 2016 · Acharya & Chellappan, Big Data and Analytics, Wiley, 2016.
WATCH OUT
The syllabus numbers its two textbooks 1 and 2, then numbers the reference books 3 and 4 rather than restarting — so the four titles read as one list and the distinction between prescribed and recommended is lost. See review finding D21.
Draw the HDFS write path once, from memory. Client → NameNode → pipeline of DataNodes → acknowledgement. If you can draw it, Unit 2 is done.
Do the block arithmetic by hand. 260 MB at a 128 MB block size is 128 + 128 + 4, not three full blocks. This is the most examined calculation in the course.
Write word count and the inverted index without looking. Map, shuffle, reduce — and be able to say what crosses the network at each step.
Learn one comparison table per unit. HDFS vs a normal filesystem; FIFO vs Fair vs Capacity; Hive vs Pig vs Spark; Avro vs Parquet; HBase vs Hive vs MongoDB. Examiners ask for these directly.
Run the labs. Fourteen of the seventeen execute, including real Spark. Numbers you have watched appear are numbers you remember.
KEY INSIGHT
The shuffle is the only step that costs real money. Map and reduce are
local and parallel; the shuffle moves every intermediate record across the
network. Combiners, partitioners, reduceByKey instead of groupByKey,
bucketed joins, broadcast joins — every optimisation in big data is an
attempt to shuffle less.
Say that in an exam and you have framed the whole subject.
What big data means and why a byte count is the wrong definition; the five Vs and which of them actually drive a design; big data against a traditional RDBMS; the Hadoop ecosystem by layer; the four core components including the one everybody forgets; the two-cluster architecture, data locality and why the cloud has weakened it; use cases, and when Hadoop is the wrong answer.
UNIT 2The NameNode and DataNode split, and why the NameNode never touches your data; block arithmetic and why a block is a maximum rather than an allocation; the small-files problem measured; replica placement, rack awareness and what replication 3 really guarantees; the read and write paths, and why a pipeline; DataNode and NameNode failure; YARN's four components and the three schedulers, measured.
UNIT 3The map, shuffle and reduce phases, and why the shuffle is the only expensive step; combiners, their measured saving and when one silently gives wrong answers; partitioning and skew; the three Java details that are examined; Hive as a compiler, partitioning, bucketing and managed against external tables; Pig as a dataflow language and its two operators with no SQL equivalent; Crunch, the abstraction ladder, and why Spark replaced MapReduce.
UNIT 4Sqoop's whole trick in two lines of SQL, split columns and what skew costs; incremental imports and the deletes they never catch; Flume's source, channel and sink, back-pressure, channel durability and the defaults that manufacture the small-files problem; Avro, Parquet and SequenceFile; schema evolution demonstrated; column projection and predicate pushdown; batch and streaming joined, the fan trap, and Lambda against Kappa.
UNIT 5Why HBase exists and what HDFS cannot do; the sparse sorted-map model, versions, tombstones and compaction; row-key design and the hotspot-against-scans trade you cannot escape; HBase against Hive and against MongoDB; CAP and why HBase is CP; ZooKeeper's ephemeral and sequential znodes, leader election, distributed locking and quorum arithmetic; Spark's RDDs, lineage, lazy evaluation, caching and the HBase integration.
PRACTICEExam-style questions with fully worked solutions.
LABEvery prescribed lab experiment, with code and expected output.