Skip to the content
On this page
  1. The one thing to understand before anything else
  2. What actually runs here, and what does not
  3. Course objectives (verbatim)
  4. The five units, and how they fit together
  5. Also here
  6. Cross-course connections
  7. Textbooks
  8. How to study this course
  9. Units in this Course

Part of the data-platform path: Big Data Technologies, Cloud Computing for Data Science, Time Series Analysis and Forecasting and Data Engineering and MLOps.


The one thing to understand before anything else

"Big data" is not a size. It is the point at which the data does not fit on one machine, and everything you know stops working.

That threshold moves. In 2006 it was a few hundred gigabytes; a laptop now holds two terabytes and a single cloud VM can be rented with 24 TB of RAM. So the useful definition is not "more than N bytes" — it is:

NOTE

Data is big when the cost of moving it exceeds the cost of computing on it.

Every design decision in this course follows from that one sentence.

The old assumption What replaces it Where you see it
Move data to the code Move the code to the data HDFS data locality, Unit 2
Disks are reliable Disks fail constantly; replicate Replication factor 3, Unit 2
One machine, scale UP Many machines, scale OUT YARN, Unit 2
Schema before data Schema when you read it Hive external tables, Unit 3
Update rows in place Write once, append only HDFS, Unit 2
A transaction is atomic Eventual consistency is enough HBase, Unit 5

⚠️ The honest framing your syllabus does not give you

This course teaches a stack that peaked around 2015. MapReduce has been superseded by Spark; Sqoop and Flume were retired to the Apache Attic in 2021; Hadoop-on-premises is losing to object storage plus a query engine.

That is not a reason to skip it, and here is why:

So learn the concepts as permanent and the tools as historical, and say so in the exam — an answer that places Hadoop in time reads as understanding rather than memorisation.


What actually runs here, and what does not

Everything runs.

The tools, on a Hadoop 3.3.6 cluster The checks, in Python
HDFS and YARN — hdfs dfs, fsck, dfsadmin, the sample jobs, the queues A MapReduce engine written out in full, with a visible shuffle
Java MapReduce — WordCount.java and InvertedIndex.java, compiled and run DuckDB for Hive-style SQL
Pig and Hive SQLite as the RDBMS for the Sqoop import
Sqoop, from MariaDB, and a Flume agent Avro via fastavro and Parquet via pyarrow — real files
HBase, a three-server ZooKeeper ensemble, and Spark reading HBase Apache Spark — a genuine SparkSession, real RDDs, a real shuffle

tools/data-science/run_bigdata_labs.py runs both halves and checks the answers each gives; lab.md shows what every file printed.

Updated October 2026: Hadoop and its ecosystem could not be installed where these notes are checked, and every tool file said NOT EXECUTED. They now install from archive.apache.org (tools/data-science/setup_hadoop.sh). Running them found what reading them had not, and twelve of the fifteen files now carry corrections, each noted in the file and in lab.md.


Course objectives (verbatim)

  1. Introduce students to the concepts, characteristics, and challenges of Big Data.

  2. Familiarize students with the Hadoop ecosystem and its core components (HDFS, YARN, MapReduce).

  3. Develop practical knowledge of distributed storage and parallel processing in Hadoop.

  4. Provide hands-on exposure to data ingestion tools (Sqoop, Flume) and serialization techniques.

  5. Enable students to explore NoSQL databases (HBase), coordination services (ZooKeeper), and HadoopSpark integration for large-scale data analysis.

WATCH OUT

⚠️ "HadoopSpark" is two words

Objective 5 and Outcome 5 both read HadoopSpark, with the space lost. They mean Hadoop and Spark — two systems, and the distinction is the whole point of Unit 5. See review finding D24.

The five units, and how they fit together

Unit Question it answers
1 What is big data, and what is in the Hadoop box?
2 Where does the data live, and who decides what runs?
3 How do you compute over it?
4 How does it get in, and in what format?
5 What if you need random access, coordination or speed?

Read them in order. Unit 2 is the load-bearing one: HDFS's design decisions explain almost everything else in the course, and a student who understands blocks, replication and the NameNode's memory can derive most of Units 3–5.


Also here

Cross-course connections

This course does not stand alone, and the labs make the links checkable.

From To What is shared
Business Intelligence Tools (BI) experiments 10, 14, 17 The same nine-row star schema, imported not copied. South = ₹10,360 is produced by DAX, DuckDB and Spark, and the suite fails if they disagree.
Database Management Systems (DBMS) Units 3 and 5 Hive is SQL that compiles to a job; HBase is what you get when you drop joins, indexes and transactions.
Document Oriented Database (MongoDB) Unit 5 Two NoSQL stores, compared directly. The sharpest difference: MongoDB has secondary indexes; HBase does not.
Python for Data Analysis and Visualization throughout pandas is the single-machine version of every operation here.
Machine Learning (ML) Unit 3 Spark's MLlib is the same algorithms, distributed — and the reason Spark beat MapReduce is iterative ML jobs.

Textbooks

References: BIG DATA, Black Book, DreamTech Press, 2016 · Acharya & Chellappan, Big Data and Analytics, Wiley, 2016.

WATCH OUT

⚠️ The reference list starts at 3

The syllabus numbers its two textbooks 1 and 2, then numbers the reference books 3 and 4 rather than restarting — so the four titles read as one list and the distinction between prescribed and recommended is lost. See review finding D21.

How to study this course

  1. Draw the HDFS write path once, from memory. Client → NameNode → pipeline of DataNodes → acknowledgement. If you can draw it, Unit 2 is done.

  2. Do the block arithmetic by hand. 260 MB at a 128 MB block size is 128 + 128 + 4, not three full blocks. This is the most examined calculation in the course.

  3. Write word count and the inverted index without looking. Map, shuffle, reduce — and be able to say what crosses the network at each step.

  4. Learn one comparison table per unit. HDFS vs a normal filesystem; FIFO vs Fair vs Capacity; Hive vs Pig vs Spark; Avro vs Parquet; HBase vs Hive vs MongoDB. Examiners ask for these directly.

  5. Run the labs. Fourteen of the seventeen execute, including real Spark. Numbers you have watched appear are numbers you remember.

KEY INSIGHT

The single highest-value fact in the course

The shuffle is the only step that costs real money. Map and reduce are local and parallel; the shuffle moves every intermediate record across the network. Combiners, partitioners, reduceByKey instead of groupByKey, bucketed joins, broadcast joins — every optimisation in big data is an attempt to shuffle less.

Say that in an exam and you have framed the whole subject.

Units in this Course

UNIT 1

Foundations of Big Data and the Hadoop Ecosystem

What big data means and why a byte count is the wrong definition; the five Vs and which of them actually drive a design; big data against a traditional RDBMS; the Hadoop ecosystem by layer; the four core components including the one everybody forgets; the two-cluster architecture, data locality and why the cloud has weakened it; use cases, and when Hadoop is the wrong answer.

UNIT 2

Hadoop Distributed File System and YARN

The NameNode and DataNode split, and why the NameNode never touches your data; block arithmetic and why a block is a maximum rather than an allocation; the small-files problem measured; replica placement, rack awareness and what replication 3 really guarantees; the read and write paths, and why a pipeline; DataNode and NameNode failure; YARN's four components and the three schedulers, measured.

UNIT 3

MapReduce and High-Level Tools

The map, shuffle and reduce phases, and why the shuffle is the only expensive step; combiners, their measured saving and when one silently gives wrong answers; partitioning and skew; the three Java details that are examined; Hive as a compiler, partitioning, bucketing and managed against external tables; Pig as a dataflow language and its two operators with no SQL equivalent; Crunch, the abstraction ladder, and why Spark replaced MapReduce.

UNIT 4

Data Ingestion and Serialization

Sqoop's whole trick in two lines of SQL, split columns and what skew costs; incremental imports and the deletes they never catch; Flume's source, channel and sink, back-pressure, channel durability and the defaults that manufacture the small-files problem; Avro, Parquet and SequenceFile; schema evolution demonstrated; column projection and predicate pushdown; batch and streaming joined, the fan trap, and Lambda against Kappa.

UNIT 5

NoSQL and Ecosystem Enhancements

Why HBase exists and what HDFS cannot do; the sparse sorted-map model, versions, tombstones and compaction; row-key design and the hotspot-against-scans trade you cannot escape; HBase against Hive and against MongoDB; CAP and why HBase is CP; ZooKeeper's ephemeral and sequential znodes, leader election, distributed locking and quorum arithmetic; Spark's RDDs, lineage, lazy evaluation, caching and the HBase integration.

PRACTICE

Practice

Exam-style questions with fully worked solutions.

LAB

Lab

Every prescribed lab experiment, with code and expected output.

Next course in learning order: Cloud Computing for Data Science Data & platforms