Skip to the content

How These Notes Are Organised

The four units are divided by the kind of supervision signal the algorithm learns from — labelled data, unlabelled data, a mixture of the two, or a reward. That is the division that actually explains why the algorithms differ, so it is the one the notes are built on.

Each unit opens on its own page with the idea behind it and its learning outcomes, then splits into topic pages. Every algorithm on a topic page is presented the same way:

Definition → Mathematical foundation → How it works → Assumptions and failure modes → Three worked examples (finance, agriculture, medicine) → Runnable Python and R code.

The Four Units

UNIT 1

🎯 Supervised Learning

Learn a mapping from labelled examples, then predict the label of data you have never seen.

10 algorithms in 2 topics • Classification • Regression
UNIT 2

🔍 Unsupervised Learning

No labels at all — find the structure that is already in the data.

6 algorithms in 3 topics • Clustering • Association Rule Learning • Anomaly Detection
UNIT 3

🔄 Semi-Supervised Learning

A few labels and a large unlabelled pool — the situation you are actually in most of the time.

3 algorithms in 2 topics • Inductive Methods • Transductive Methods
UNIT 4

🕹️ Reinforcement Learning

No dataset at all — an agent, an environment, and a reward signal to learn from.

4 algorithms in 2 topics • Model-Free Methods • Model-Based Methods
REFERENCE

📜 Scope and Coverage

Every algorithm in one table, what each unit assumes you already know, and an honest list of what these notes do not yet cover.

Full inventory • prerequisites • known gaps

All Topic Pages

UnitTopicAlgorithms
Supervised Learning1.1 ClassificationNaive Bayes Classifier, Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree (Classification)
1.2 RegressionLinear Regression, Polynomial Regression, Ridge Regression (L2 Regularisation), Lasso Regression (L1 Regularisation), Random Forest Regression
Unsupervised Learning2.1 ClusteringK-Means Clustering, DBSCAN, Hierarchical / Agglomerative Clustering
2.2 Association Rule LearningApriori Algorithm, FP-Growth Algorithm
2.3 Anomaly DetectionIsolation Forest
Semi-Supervised Learning3.1 Inductive MethodsSelf-Training, Co-Training
3.2 Transductive MethodsLabel Propagation
Reinforcement Learning4.1 Model-Free MethodsQ-Learning, Policy Gradient (REINFORCE)
4.2 Model-Based MethodsDyna-Q, Value Iteration (Dynamic Programming)

Running the Code

Every Python pane is self-contained and simulates its own data — there is nothing to download. All 23 panes are verified to run end to end.

pip install numpy pandas scikit-learn mlxtend

The R panes use e1071, caret, class, rpart, randomForest, glmnet, cluster, dbscan and arules, depending on the algorithm.

About the figures in the examples. The worked examples describe realistic settings, but the code simulates its data, so any accuracy a pane prints is a property of that simulation, not a published result. Where an example uses a real public dataset it is named, so the number can be checked.