Supervised learning is the setting most people mean when they say “machine learning”. You are given pairs \((\mathbf{x}_i, y_i)\) — features and the answer — and the task is to learn a function \(f\) such that \(f(\mathbf{x}) \approx y\) on data drawn from the same distribution. Everything else in this unit follows from two questions: what shape is \(y\), and what shape do you allow \(f\) to take?
By the end of this unit you should be able to:
Predicting which class a point belongs to: probabilistic, linear, instance-based, margin-based and tree-based approaches.
5 algorithms • Naive Bayes Classifier, Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree (Classification) 1.2Predicting a number: least squares, basis expansion, the two regularisation penalties, and an ensemble of trees.
5 algorithms • Linear Regression, Polynomial Regression, Ridge Regression (L2 Regularisation), Lasso Regression (L1 Regularisation), Random Forest Regression| No. | Algorithm | Character | Topic |
|---|---|---|---|
| 1.1 | Naive Bayes Classifier | Probabilistic • Generative | 1.1 Classification |
| 1.2 | Logistic Regression | Linear • Discriminative | 1.1 Classification |
| 1.3 | K-Nearest Neighbors (KNN) | Instance-Based • Non-Parametric | 1.1 Classification |
| 1.4 | Support Vector Machine (SVM) | Margin Maximisation • Kernel Methods | 1.1 Classification |
| 1.5 | Decision Tree (Classification) | Tree-Based • Interpretable | 1.1 Classification |
| 1.6 | Linear Regression | Parametric • Closed-Form | 1.2 Regression |
| 1.7 | Polynomial Regression | Non-Linear • Feature Engineering | 1.2 Regression |
| 1.8 | Ridge Regression (L2 Regularisation) | Regularisation • Shrinkage | 1.2 Regression |
| 1.9 | Lasso Regression (L1 Regularisation) | Regularisation • Feature Selection | 1.2 Regression |
| 1.10 | Random Forest Regression | Ensemble • Bagging | 1.2 Regression |