Skip to the content
On this page
  1. What this course actually is
  2. Where it sits in the degree
  3. Course objectives (verbatim)
  4. Units in this Course

What this course actually is

Data mining is finding patterns in data that nobody put there deliberately. A supermarket's till system was built to take payments; the fact that nappies and beer sell together on Friday evenings is a pattern hiding in the exhaust of that system. Data Mining is the set of algorithms for extracting such patterns at scale.

It is the most algorithmic course in the catalogue. Where Data Science with R taught you to call lm() and Python for Data Analysis and Visualization will teach df.groupby(), this course asks you to trace Apriori by hand, compute an information gain, and run K-Means to convergence on paper. That is deliberate, and it is what the exam tests. An algorithm you can trace by hand is one you can debug when a library gives you a wrong answer.

Where it sits in the degree

From You already have Used here for
Statistical Foundations for Data Science Mean, variance, probability, Bayes Naïve Bayes, evaluation metrics
Database Management Systems SQL, schemas, normalisation Unit 1's star schema — deliberately denormalised
Data Science with R K-Means, TF-IDF, confusion matrices Revisited properly, with the arithmetic
Python for Data Analysis and Visualization (parallel) NumPy and Pandas The lab equivalents are written in them

Unit 1 is the surprise: it is not mining at all but data warehousing, and it exists because you cannot mine what you cannot assemble. It also directly contradicts Database Management Systems' normalisation teaching, on purpose — §1.6 explains why.

Course objectives (verbatim)

  1. Provide an understanding of data warehousing concepts, architecture, and OLAP operations for effective storage, modeling, and analysis.

  2. Develop knowledge of data mining fundamentals, tasks, and preprocessing techniques to prepare data for mining.

  3. Introduce students to association rule mining algorithms for discovering hidden patterns and relationships in large datasets.

  4. Enable learners to apply classification techniques (decision trees, Bayesian, nearest neighbor, rule-based) for predictive modeling.

  5. Equip students with knowledge of clustering paradigms and algorithms (partitioning, hierarchical, density-based, categorical) for data grouping and pattern discovery.

Units in this Course

UNIT 1

Data Warehousing and OLAP

Inmon’s four characteristics; OLTP versus OLAP; three-tier architecture and ETL; the multidimensional model; fact and dimension tables; star, snowflake and fact constellation schemas; the cube and the five OLAP operations.

UNIT 2

Data Mining and Preprocessing

Definitions and the KDD process; predictive versus descriptive tasks; cleaning, missing data and noise; the curse of dimensionality and PCA; feature subset selection; discretization and binarization; normalisation; similarity and dissimilarity measures; issues, ethics and applications.

UNIT 3

Association Analysis

Support, confidence and lift; why confidence alone misleads; the Apriori principle and algorithm; rule generation; Partition, Pincer-Search and Dynamic Itemset Counting; the FP-tree and FP-Growth; generalized rules and item constraints.

UNIT 4

Classification

Decision trees and the best split; entropy, information gain, gain ratio and Gini; ID3, C4.5 and CART; overfitting and pruning; the confusion matrix, precision, recall, F1, ROC and AUC; rule-based classifiers; k-nearest neighbour; Naïve Bayes and Laplace smoothing.

UNIT 5

Clustering Techniques

Clustering paradigms and validity measures; K-Means and its five weaknesses; K-Medoids and PAM; agglomerative clustering and linkage criteria; DBSCAN; BIRCH and the CF-tree; categorical clustering with STIRR, ROCK and CACTUS.

PRACTICE

Practice

Exam-style questions with fully worked solutions.

LAB

Lab

Every prescribed lab experiment, with code and expected output.

Next course in learning order: Machine Learning Machine learning & AI