The four units are divided by the kind of supervision signal the algorithm learns from — labelled data, unlabelled data, a mixture of the two, or a reward. That is the division that actually explains why the algorithms differ, so it is the one the notes are built on.
Each unit opens on its own page with the idea behind it and its learning outcomes, then splits into topic pages. Every algorithm on a topic page is presented the same way:
Learn a mapping from labelled examples, then predict the label of data you have never seen.
10 algorithms in 2 topics • Classification • Regression UNIT 2No labels at all — find the structure that is already in the data.
6 algorithms in 3 topics • Clustering • Association Rule Learning • Anomaly Detection UNIT 3A few labels and a large unlabelled pool — the situation you are actually in most of the time.
3 algorithms in 2 topics • Inductive Methods • Transductive Methods UNIT 4No dataset at all — an agent, an environment, and a reward signal to learn from.
4 algorithms in 2 topics • Model-Free Methods • Model-Based Methods REFERENCEEvery algorithm in one table, what each unit assumes you already know, and an honest list of what these notes do not yet cover.
Full inventory • prerequisites • known gaps| Unit | Topic | Algorithms |
|---|---|---|
| Supervised Learning | 1.1 Classification | Naive Bayes Classifier, Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree (Classification) |
| 1.2 Regression | Linear Regression, Polynomial Regression, Ridge Regression (L2 Regularisation), Lasso Regression (L1 Regularisation), Random Forest Regression | |
| Unsupervised Learning | 2.1 Clustering | K-Means Clustering, DBSCAN, Hierarchical / Agglomerative Clustering |
| 2.2 Association Rule Learning | Apriori Algorithm, FP-Growth Algorithm | |
| 2.3 Anomaly Detection | Isolation Forest | |
| Semi-Supervised Learning | 3.1 Inductive Methods | Self-Training, Co-Training |
| 3.2 Transductive Methods | Label Propagation | |
| Reinforcement Learning | 4.1 Model-Free Methods | Q-Learning, Policy Gradient (REINFORCE) |
| 4.2 Model-Based Methods | Dyna-Q, Value Iteration (Dynamic Programming) |
Every Python pane is self-contained and simulates its own data — there is nothing to download. All 23 panes are verified to run end to end.
pip install numpy pandas scikit-learn mlxtend
The R panes use e1071, caret, class, rpart,
randomForest, glmnet, cluster, dbscan and
arules, depending on the algorithm.