Skip to the content

What This Unit Is About

Here there is no \(y\). You are given \(\mathbf{x}_1, \dots, \mathbf{x}_N\) and asked what structure they contain: which points group together, which items co-occur, which points do not belong. Because there is no ground truth to score against, the hard part of unsupervised learning is rarely the algorithm — it is deciding whether the structure you found is real.

Learning Outcomes

By the end of this unit you should be able to:

  1. Choose a clustering algorithm from the shape of the clusters you expect, and justify the choice.
  2. Select \(K\) for K-Means and \(\varepsilon\) for DBSCAN using the elbow and k-distance plots rather than by guessing.
  3. Compute support, confidence and lift for an association rule, and explain why high confidence alone is not evidence.
  4. Explain how FP-Growth avoids Apriori’s candidate generation, and when that advantage actually matters.
  5. Score anomalies by isolation path length, and state why the contamination rate is the assumption that governs everything.

Topics in This Unit

2.1

🔵 Clustering

Partitioning points into groups: by centroid, by density, and by a hierarchy you can cut at any height.

3 algorithms • K-Means Clustering, DBSCAN, Hierarchical / Agglomerative Clustering
2.2

🔗 Association Rule Learning

Which items co-occur more often than chance would predict — and how to mine them without enumerating every subset.

2 algorithms • Apriori Algorithm, FP-Growth Algorithm
2.3

🚨 Anomaly Detection

Finding the points that do not belong, when rarity itself is the only definition you have.

1 algorithm • Isolation Forest

Every Algorithm in Unsupervised Learning

No.AlgorithmCharacterTopic
2.1K-Means ClusteringCentroid-Based • Partitional2.1 Clustering
2.2DBSCANDensity-Based • Arbitrary Shapes2.1 Clustering
2.3Hierarchical / Agglomerative ClusteringHierarchical • Dendrogram2.1 Clustering
2.4Apriori AlgorithmFrequent Itemsets • Association Rules2.2 Association Rule Learning
2.5FP-Growth AlgorithmTree-Based Mining • Scalable2.2 Association Rule Learning
2.6Isolation ForestAnomaly Detection • Tree Ensemble2.3 Anomaly Detection