Data mining is finding patterns in data that nobody put there deliberately. A supermarket's till system was built to take payments; the fact that nappies and beer sell together on Friday evenings is a pattern hiding in the exhaust of that system. Data Mining is the set of algorithms for extracting such patterns at scale.
It is the most algorithmic course in the catalogue. Where Data Science with R taught
you to call lm() and Python for Data Analysis and Visualization will teach df.groupby(), this course asks
you to trace Apriori by hand, compute an information gain, and run K-Means to
convergence on paper. That is deliberate, and it is what the exam tests.
An algorithm you can trace by hand is one you can debug when a library gives
you a wrong answer.
| From | You already have | Used here for |
|---|---|---|
| Statistical Foundations for Data Science | Mean, variance, probability, Bayes | Naïve Bayes, evaluation metrics |
| Database Management Systems | SQL, schemas, normalisation | Unit 1's star schema — deliberately denormalised |
| Data Science with R | K-Means, TF-IDF, confusion matrices | Revisited properly, with the arithmetic |
| Python for Data Analysis and Visualization (parallel) | NumPy and Pandas | The lab equivalents are written in them |
Unit 1 is the surprise: it is not mining at all but data warehousing, and it exists because you cannot mine what you cannot assemble. It also directly contradicts Database Management Systems' normalisation teaching, on purpose — §1.6 explains why.
Provide an understanding of data warehousing concepts, architecture, and OLAP operations for effective storage, modeling, and analysis.
Develop knowledge of data mining fundamentals, tasks, and preprocessing techniques to prepare data for mining.
Introduce students to association rule mining algorithms for discovering hidden patterns and relationships in large datasets.
Enable learners to apply classification techniques (decision trees, Bayesian, nearest neighbor, rule-based) for predictive modeling.
Equip students with knowledge of clustering paradigms and algorithms (partitioning, hierarchical, density-based, categorical) for data grouping and pattern discovery.
Inmon’s four characteristics; OLTP versus OLAP; three-tier architecture and ETL; the multidimensional model; fact and dimension tables; star, snowflake and fact constellation schemas; the cube and the five OLAP operations.
UNIT 2Definitions and the KDD process; predictive versus descriptive tasks; cleaning, missing data and noise; the curse of dimensionality and PCA; feature subset selection; discretization and binarization; normalisation; similarity and dissimilarity measures; issues, ethics and applications.
UNIT 3Support, confidence and lift; why confidence alone misleads; the Apriori principle and algorithm; rule generation; Partition, Pincer-Search and Dynamic Itemset Counting; the FP-tree and FP-Growth; generalized rules and item constraints.
UNIT 4Decision trees and the best split; entropy, information gain, gain ratio and Gini; ID3, C4.5 and CART; overfitting and pruning; the confusion matrix, precision, recall, F1, ROC and AUC; rule-based classifiers; k-nearest neighbour; Naïve Bayes and Laplace smoothing.
UNIT 5Clustering paradigms and validity measures; K-Means and its five weaknesses; K-Medoids and PAM; agglomerative clustering and linkage criteria; DBSCAN; BIRCH and the CF-tree; categorical clustering with STIRR, ROCK and CACTUS.
PRACTICEExam-style questions with fully worked solutions.
LABEvery prescribed lab experiment, with code and expected output.