Skip to the content

Algorithms on This Page

1.1 Naive Bayes Classifier 1.2 Logistic Regression 1.3 K-Nearest Neighbors (KNN) 1.4 Support Vector Machine (SVM) 1.5 Decision Tree (Classification)

1.1  Naive Bayes Classifier

ProbabilisticGenerative
DEFINITION

Naive Bayes is a probabilistic classifier based on Bayes' theorem that assumes conditional independence between every pair of features given the class label. Despite this "naive" assumption, it performs remarkably well in many real-world scenarios.

1.1.1  Mathematical Foundation

FORMULAE

Bayes' Theorem:

\[P(C_k \mid \mathbf{x}) = \frac{P(\mathbf{x} \mid C_k)\, P(C_k)}{P(\mathbf{x})}\]

Naive Independence Assumption:

\[P(\mathbf{x} \mid C_k) = \prod_{i=1}^{n} P(x_i \mid C_k)\]

Classification Rule:

\[\hat{y} = \arg\max_{k} \left[ \log P(C_k) + \sum_{i=1}^{n} \log P(x_i \mid C_k) \right]\]

Gaussian Naive Bayes (for continuous features):

\[P(x_i \mid C_k) = \frac{1}{\sqrt{2\pi\sigma_{ik}^2}} \exp\!\left(-\frac{(x_i - \mu_{ik})^2}{2\sigma_{ik}^2}\right)\]

1.1.2  How It Works

The algorithm estimates the prior probability \(P(C_k)\) from class frequencies and the likelihood \(P(x_i|C_k)\) from training data (Gaussian distribution for continuous features, multinomial for text). At prediction time, it picks the class that maximises the posterior probability. The log-sum trick prevents numerical underflow when multiplying many small probabilities. Despite its simplicity, Naive Bayes trains extremely fast and generalises well even with small datasets.

1.1.3  Assumptions and Failure Modes

ASSUMES
  • Features conditionally independent given the class
  • Correct likelihood family (Gaussian / multinomial)
BREAKS WHEN
  • Features are strongly correlated — posteriors become badly over-confident
  • An unseen category gives zero likelihood without Laplace smoothing

1.1.4  Worked Examples

FINANCE

🏦 Loan Default Prediction

A bank classifies loan applicants as default or no default using credit score, income, and debt ratio. With 2,000 labelled historical accounts, Gaussian Naive Bayes is fast to fit and easy to explain to a credit committee. How accurate it is depends almost entirely on the base default rate and on how far apart the credit-score distributions sit — run the code below to see what it reaches on simulated data of this shape.

CreditScoreIncome ($k)DebtRatioDefault
720850.28No
540420.71Yes
680630.45No
AGRICULTURE

🌾 Crop Disease Detection

Using leaf features (colour index, lesion size, humidity), farmers classify crop disease type: rust, blight, or healthy. Field sensor data from 500 plots trains the model.

ColourIdxLesionMMHumidity%Class
0.823.178Rust
0.310.255Healthy
0.678.485Blight
MEDICINE

💊 Tumour Diagnosis

Pathologists use cell-nucleus features (radius, texture, perimeter) to classify tumours as malignant or benign. This is the setting of the public Wisconsin Diagnostic Breast Cancer dataset — 569 samples, 30 features — on which Gaussian Naive Bayes reaches roughly 93% accuracy. Reproduce it with sklearn.datasets.load_breast_cancer().

RadiusTexturePerimeterDiagnosis
17.9910.38122.8Malignant
13.5414.3687.46Benign

1.1.5  Code

Naive Bayes
import numpy as np
import pandas as pd
from sklearn.naive_bayes import GaussianNB
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report, confusion_matrix
from sklearn.preprocessing import StandardScaler

np.random.seed(42)
n = 600

# ── Simulate: Financial Loan Default Dataset ──────────────────
no_default = pd.DataFrame({
    'credit_score': np.random.normal(700, 50, 420),
    'income_k':     np.random.normal(75, 18, 420),
    'debt_ratio':   np.random.normal(0.28, 0.08, 420),
    'default':      np.zeros(420, dtype=int)
})
default = pd.DataFrame({
    'credit_score': np.random.normal(530, 60, 180),
    'income_k':     np.random.normal(38, 12, 180),
    'debt_ratio':   np.random.normal(0.72, 0.12, 180),
    'default':      np.ones(180, dtype=int)
})
df = pd.concat([no_default, default]).sample(frac=1, random_state=42).reset_index(drop=True)

X = df[['credit_score','income_k','debt_ratio']].values
y = df['default'].values

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=42)

# ── Train Gaussian Naive Bayes ────────────────────────────────
model = GaussianNB()
model.fit(X_train, y_train)

# ── Evaluate ──────────────────────────────────────────────────
y_pred = model.predict(X_test)
print("=== Naive Bayes — Loan Default Prediction ===")
print(classification_report(y_test, y_pred, target_names=['No Default','Default']))
print("Confusion Matrix:\n", confusion_matrix(y_test, y_pred))

# ── Posterior probabilities for new applicant ─────────────────
new_applicant = np.array([[650, 55, 0.45]])
proba = model.predict_proba(new_applicant)
print(f"\nNew applicant — P(no default)={proba[0,0]:.3f}, P(default)={proba[0,1]:.3f}")

# ─── Agriculture: Crop Disease ─────────────────────────────────
np.random.seed(7)
classes = ['Healthy','Rust','Blight']
crop_data = pd.DataFrame({
    'colour_idx': np.concatenate([np.random.normal(0.3,0.07,200),
                                   np.random.normal(0.82,0.05,200),
                                   np.random.normal(0.65,0.08,200)]),
    'lesion_mm':  np.concatenate([np.random.normal(0.3,0.1,200),
                                   np.random.normal(3.2,0.6,200),
                                   np.random.normal(8.5,1.2,200)]),
    'humidity':   np.concatenate([np.random.normal(52,6,200),
                                   np.random.normal(78,5,200),
                                   np.random.normal(86,4,200)]),
    'disease':    [0]*200 + [1]*200 + [2]*200
})
Xc = crop_data[['colour_idx','lesion_mm','humidity']].values
yc = crop_data['disease'].values
Xc_tr, Xc_te, yc_tr, yc_te = train_test_split(Xc, yc, test_size=0.2, random_state=42)
nb_crop = GaussianNB().fit(Xc_tr, yc_tr)
print(f"\nCrop Disease Accuracy: {nb_crop.score(Xc_te, yc_te):.3f}")
library(e1071)
library(caret)
set.seed(42)

# ── Simulate Loan Default Dataset ─────────────────────────────
n_good <- 420; n_bad <- 180
good <- data.frame(
  credit_score = rnorm(n_good, 700, 50),
  income_k     = rnorm(n_good, 75, 18),
  debt_ratio   = rnorm(n_good, 0.28, 0.08),
  default      = factor(rep("No", n_good))
)
bad <- data.frame(
  credit_score = rnorm(n_bad, 530, 60),
  income_k     = rnorm(n_bad, 38, 12),
  debt_ratio   = rnorm(n_bad, 0.72, 0.12),
  default      = factor(rep("Yes", n_bad))
)
df <- rbind(good, bad)[sample(600), ]

# ── Train/Test Split ──────────────────────────────────────────
idx <- createDataPartition(df$default, p = 0.75, list = FALSE)
train_df <- df[idx, ]; test_df <- df[-idx, ]

# ── Train Naive Bayes ─────────────────────────────────────────
nb_model <- naiveBayes(default ~ ., data = train_df)

# ── Evaluate ──────────────────────────────────────────────────
preds <- predict(nb_model, test_df)
cm <- confusionMatrix(preds, test_df$default)
print(cm)

# ── Posterior probability for new applicant ───────────────────
new_app <- data.frame(credit_score=650, income_k=55, debt_ratio=0.45)
proba   <- predict(nb_model, new_app, type="raw")
cat("\nNew Applicant Posterior Probabilities:\n")
print(proba)

# ─── Agriculture: Crop Disease ─────────────────────────────────
set.seed(7)
crop <- data.frame(
  colour_idx = c(rnorm(200,0.3,0.07), rnorm(200,0.82,0.05), rnorm(200,0.65,0.08)),
  lesion_mm  = c(rnorm(200,0.3,0.1),  rnorm(200,3.2,0.6),   rnorm(200,8.5,1.2)),
  humidity   = c(rnorm(200,52,6),     rnorm(200,78,5),      rnorm(200,86,4)),
  disease    = factor(rep(c("Healthy","Rust","Blight"), each=200))
)
nb_crop <- naiveBayes(disease ~ ., data = crop)
pred_crop <- predict(nb_crop, crop)
acc <- mean(pred_crop == crop$disease)
cat(sprintf("\nCrop Disease Accuracy: %.3f\n", acc))

1.2  Logistic Regression

LinearDiscriminative
DEFINITION

Logistic Regression is a linear classifier that models the probability of a binary (or multi-class) outcome using the sigmoid (logistic) function applied to a linear combination of features. It is trained by maximising log-likelihood via gradient descent or Newton's method.

1.2.1  Mathematical Foundation

FORMULAE

Sigmoid (logistic) function:

\[\sigma(z) = \frac{1}{1 + e^{-z}}, \quad z = \mathbf{w}^\top \mathbf{x} + b\]

Predicted probability:

\[\hat{p} = P(y=1 \mid \mathbf{x}) = \sigma(\mathbf{w}^\top \mathbf{x} + b)\]

Binary Cross-Entropy Loss:

\[\mathcal{L}(\mathbf{w}) = -\frac{1}{N}\sum_{i=1}^{N}\bigl[y_i \log\hat{p}_i + (1-y_i)\log(1-\hat{p}_i)\bigr]\]

Gradient update:

\[\mathbf{w} \leftarrow \mathbf{w} - \alpha\,\nabla_{\mathbf{w}}\mathcal{L}\]

Decision boundary: \(\hat{y} = 1\) if \(\hat{p} \ge 0.5\)

1.2.2  How It Works

Logistic regression learns a linear boundary in feature space by squashing the linear combination through a sigmoid to produce a probability in \((0,1)\). The loss penalises confident wrong predictions heavily. Regularisation terms (L1 or L2) prevent overfitting. For multi-class problems, the softmax function generalises this to \(K\) classes: \(P(y=k|\mathbf{x}) = e^{\mathbf{w}_k^\top\mathbf{x}}/\sum_j e^{\mathbf{w}_j^\top\mathbf{x}}\). Coefficients are interpretable as log-odds ratios.

1.2.3  Assumptions and Failure Modes

ASSUMES
  • Log-odds linear in the features
  • Observations independent
BREAKS WHEN
  • Classes are perfectly separable — coefficients diverge without regularisation
  • The true boundary is curved and no interaction terms are added

1.2.4  Worked Examples

FINANCE

🏦 Credit Card Fraud Detection

Using transaction amount, frequency, and location deviation, banks predict fraud probability per transaction. The code below simulates 10,000 transactions at a 3% fraud rate and prints the AUC-ROC it actually achieves. Treat that number as a property of the simulation: real card fraud is far less separable, and class imbalance makes AUC-ROC a kinder metric than precision at a fixed alert budget.

Amount ($)Freq/dayLocDevFraud
1,24018850kmYes
4532kmNo
3,200221,400kmYes
AGRICULTURE

🌿 Irrigation Need Prediction

Using soil moisture, temperature, and days since last rain, agronomists predict whether a field needs irrigation today (1) or not (0). 400 daily sensor records used for training.

Moisture%Temp°CDaysDryIrrigate
18347Yes
52221No
24315Yes
MEDICINE

🩺 Diabetes Onset Prediction

Using glucose level, BMI, age, and blood pressure, clinicians predict diabetes onset within 5 years. This is the setting of the public Pima Indians Diabetes dataset — 768 records, 8 features — on which logistic regression reaches roughly 77% accuracy, with glucose carrying the largest standardised coefficient. The dataset also carries well-documented missing values coded as zeros, which is itself worth teaching.

GlucoseBMIAgeDiabetic
14833.650Yes
8526.631No
18323.332Yes

1.2.5  Code

Logistic Regression
import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.metrics import roc_auc_score, classification_report
from sklearn.preprocessing import StandardScaler

np.random.seed(0)

# ── Simulate Diabetes Dataset ─────────────────────────────────
n_pos, n_neg = 268, 500
diabetic = pd.DataFrame({
    'glucose':   np.random.normal(141, 31, n_pos),
    'bmi':       np.random.normal(35.4, 7.1, n_pos),
    'age':       np.random.normal(48, 10, n_pos),
    'bp':        np.random.normal(76, 12, n_pos),
    'label':     np.ones(n_pos)
})
healthy = pd.DataFrame({
    'glucose':   np.random.normal(110, 26, n_neg),
    'bmi':       np.random.normal(30.3, 7.9, n_neg),
    'age':       np.random.normal(31, 11, n_neg),
    'bp':        np.random.normal(70, 13, n_neg),
    'label':     np.zeros(n_neg)
})
df = pd.concat([diabetic, healthy]).sample(frac=1, random_state=0).reset_index(drop=True)

X = df[['glucose','bmi','age','bp']].values
y = df['label'].values

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0, stratify=y)
scaler = StandardScaler()
X_train_sc = scaler.fit_transform(X_train)
X_test_sc  = scaler.transform(X_test)

# ── Train Logistic Regression ─────────────────────────────────
lr = LogisticRegression(C=1.0, max_iter=1000, random_state=0)
lr.fit(X_train_sc, y_train)

y_pred  = lr.predict(X_test_sc)
y_proba = lr.predict_proba(X_test_sc)[:,1]
print("=== Logistic Regression — Diabetes Prediction ===")
print(classification_report(y_test, y_pred, target_names=['Healthy','Diabetic']))
print(f"AUC-ROC: {roc_auc_score(y_test, y_proba):.4f}")

# ── Coefficients (log-odds) ───────────────────────────────────
features = ['Glucose','BMI','Age','BloodPressure']
for f, c in zip(features, lr.coef_[0]):
    print(f"  {f:18s}: log-odds coef = {c:+.4f}")

# ── 5-fold Cross-validation ───────────────────────────────────
cv_scores = cross_val_score(lr, scaler.fit_transform(X), y, cv=5, scoring='roc_auc')
print(f"\n5-Fold CV AUC: {cv_scores.mean():.4f} ± {cv_scores.std():.4f}")

# ── Fraud Detection (Financial) ───────────────────────────────
np.random.seed(99)
n_ok, n_fr = 9700, 300
legit = np.column_stack([np.random.normal(80,40,n_ok),
                          np.random.randint(1,6,n_ok).astype(float),
                          np.random.normal(5,20,n_ok)])
fraud = np.column_stack([np.random.normal(1200,600,n_fr),
                          np.random.randint(10,25,n_fr).astype(float),
                          np.random.normal(900,300,n_fr)])
Xf = np.vstack([legit, fraud])
yf = np.array([0]*n_ok + [1]*n_fr)
Xf_tr, Xf_te, yf_tr, yf_te = train_test_split(Xf, yf, test_size=0.2, stratify=yf, random_state=1)
sc2 = StandardScaler()
lr_fraud = LogisticRegression(class_weight='balanced', max_iter=500)
lr_fraud.fit(sc2.fit_transform(Xf_tr), yf_tr)
auc_f = roc_auc_score(yf_te, lr_fraud.predict_proba(sc2.transform(Xf_te))[:,1])
print(f"\nFraud Detection AUC-ROC: {auc_f:.4f}")
set.seed(0)
library(caret)

# ── Simulate Diabetes Dataset ─────────────────────────────────
diabetic <- data.frame(
  glucose = rnorm(268, 141, 31), bmi = rnorm(268, 35.4, 7.1),
  age     = rnorm(268, 48, 10),  bp  = rnorm(268, 76, 12),
  label   = factor(rep("Diabetic", 268))
)
healthy <- data.frame(
  glucose = rnorm(500, 110, 26), bmi = rnorm(500, 30.3, 7.9),
  age     = rnorm(500, 31, 11),  bp  = rnorm(500, 70, 13),
  label   = factor(rep("Healthy", 500))
)
df <- rbind(diabetic, healthy)[sample(768), ]

# ── Train / Test Split ────────────────────────────────────────
idx     <- createDataPartition(df$label, p=0.8, list=FALSE)
train   <- df[idx,]; test <- df[-idx,]

# ── Logistic Regression via glm ───────────────────────────────
lr_model <- glm(label ~ ., data=train, family=binomial())
summary(lr_model)

# ── Predictions ───────────────────────────────────────────────
probs <- predict(lr_model, test, type="response")
preds <- factor(ifelse(probs > 0.5, "Healthy", "Diabetic"))
cm    <- confusionMatrix(preds, test$label)
print(cm)

# ── ROC-AUC ───────────────────────────────────────────────────
if (!requireNamespace("pROC", quietly=TRUE)) install.packages("pROC")
library(pROC)
roc_obj <- roc(test$label, probs)
cat(sprintf("AUC-ROC: %.4f\n", auc(roc_obj)))

# ── Log-odds interpretation ───────────────────────────────────
cat("\nCoefficients (log-odds):\n")
print(coef(lr_model))

# ── 5-fold cross-validation ───────────────────────────────────
ctrl  <- trainControl(method="cv", number=5, classProbs=TRUE, summaryFunction=twoClassSummary)
df$label <- relevel(df$label, ref="Diabetic")
cv_lr <- train(label ~ ., data=df, method="glm", family="binomial",
               trControl=ctrl, metric="ROC",
               preProc=c("center","scale"))
cat(sprintf("CV AUC: %.4f\n", max(cv_lr$results$ROC)))

1.3  K-Nearest Neighbors (KNN)

Instance-BasedNon-Parametric
DEFINITION

KNN is a non-parametric, lazy learning algorithm that classifies a new data point by finding the \(k\) most similar training examples in feature space and taking the majority vote (classification) or average (regression) of their labels. It stores the entire training set and defers computation to query time.

1.3.1  Mathematical Foundation

FORMULAE

Euclidean Distance (most common):

\[d(\mathbf{p},\mathbf{q}) = \sqrt{\sum_{i=1}^{n}(p_i - q_i)^2}\]

Minkowski Distance (generalisation):

\[d_m(\mathbf{p},\mathbf{q}) = \left(\sum_{i=1}^{n}|p_i - q_i|^m\right)^{1/m}\]

Classification rule (majority vote):

\[\hat{y} = \arg\max_{c} \sum_{i \in \mathcal{N}_k(\mathbf{x})} \mathbf{1}[y_i = c]\]

Weighted KNN:

\[\hat{y} = \arg\max_{c} \sum_{i \in \mathcal{N}_k(\mathbf{x})} \frac{1}{d(\mathbf{x},\mathbf{x}_i)^2}\,\mathbf{1}[y_i = c]\]

The weight is undefined when the query coincides with a training point (\(d = 0\)). Guard it: return that point's label outright, or use \(1/(d + \varepsilon)^2\). scikit-learn's weights='distance' takes the first route.

1.3.2  How It Works

KNN's key hyperparameter is \(k\) — small \(k\) causes high variance (overfitting), large \(k\) causes high bias (underfitting). Feature scaling is critical because Euclidean distance is sensitive to feature magnitude. The optimal \(k\) is typically found via cross-validation. KNN has \(O(n \cdot d)\) prediction time per query, making it slow for large datasets. Approximate nearest-neighbour structures (KD-trees, ball-trees) speed up retrieval significantly.

1.3.3  Assumptions and Failure Modes

ASSUMES
  • Nearby points share labels
  • All features matter equally once scaled
BREAKS WHEN
  • Features are left unscaled — the largest-range feature dominates the distance
  • Dimensionality is high — distances concentrate and neighbours stop being near

1.3.4  Worked Examples

FINANCE

💳 Customer Product Recommendation

A bank finds the 5 most similar customers by age, balance, and transaction history to recommend products. New customer is matched to nearest neighbours who purchased investment funds.

AgeBalance $kTxnFreqProduct
3512015Invest
623408Bond
281822Card
AGRICULTURE

🌱 Soil Type Classification

Soil samples classified into sandy, loamy, or clay using pH, nitrogen, and phosphorus content. The code below simulates 500 such samples and reports the accuracy it reaches. Note what does the work: the three features have wildly different ranges, so standardising before computing distances matters more here than the choice of \(k\).

pHN (mg/kg)P (mg/kg)Soil Type
6.218245Loamy
7.89028Sandy
5.524062Clay
MEDICINE

🧬 Drug Response Classification

Patients classified into responders or non-responders based on genetic markers (SNP count, gene expression) before prescribing an antidepressant. The goal is to shorten the trial-and-error phase of prescribing; reported benefits vary widely by drug class and cohort, and pharmacogenomic panels remain an active research area rather than settled clinical practice.

SNP CountGene ExprAgeResponse
122.444Good
30.858Poor
81.935Good

1.3.5  Code

K-Nearest Neighbors
import numpy as np
import pandas as pd
from sklearn.neighbors import KNeighborsClassifier
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report

np.random.seed(42)

# ── Simulate Soil Classification Dataset ──────────────────────
def sim_soil(n, ph_mu, n_mu, p_mu, label):
    return pd.DataFrame({
        'ph': np.random.normal(ph_mu, 0.3, n),
        'nitrogen': np.random.normal(n_mu, 20, n),
        'phosphorus': np.random.normal(p_mu, 8, n),
        'soil_type': [label]*n
    })

df = pd.concat([
    sim_soil(200, 6.2, 180, 45, 'Loamy'),
    sim_soil(150, 7.8, 90,  28, 'Sandy'),
    sim_soil(150, 5.5, 240, 62, 'Clay')
]).sample(frac=1, random_state=42).reset_index(drop=True)

X = df[['ph','nitrogen','phosphorus']].values
y = df['soil_type'].values

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)

scaler = StandardScaler()
X_tr_sc = scaler.fit_transform(X_tr)
X_te_sc  = scaler.transform(X_te)

# ── Find optimal k via Grid Search ────────────────────────────
param_grid = {'n_neighbors': range(1, 21), 'weights': ['uniform','distance']}
gs = GridSearchCV(KNeighborsClassifier(), param_grid, cv=5, scoring='accuracy', n_jobs=-1)
gs.fit(X_tr_sc, y_tr)
print(f"Best k={gs.best_params_['n_neighbors']}, weight={gs.best_params_['weights']}")

best_knn = gs.best_estimator_
y_pred = best_knn.predict(X_te_sc)
print("\n=== KNN — Soil Type Classification ===")
print(classification_report(y_te, y_pred))

# ── Inspect nearest neighbours for a new sample ───────────────
new_sample = scaler.transform([[6.0, 175, 42]])
distances, indices = best_knn.kneighbors(new_sample)
print(f"\nNearest {gs.best_params_['n_neighbors']} neighbours:")
neighbour_df = df.iloc[indices[0]][['ph','nitrogen','phosphorus','soil_type']].copy()
neighbour_df['distance'] = distances[0]
print(neighbour_df.to_string(index=False))
print(f"\nPredicted soil type: {best_knn.predict(new_sample)[0]}")
library(class); library(caret); set.seed(42)

# ── Simulate Soil Dataset ─────────────────────────────────────
sim_soil <- function(n, ph_mu, n_mu, p_mu, label) {
  data.frame(ph=rnorm(n,ph_mu,.3), nitrogen=rnorm(n,n_mu,20),
             phosphorus=rnorm(n,p_mu,8), soil_type=factor(rep(label,n)))
}
df <- rbind(sim_soil(200,6.2,180,45,"Loamy"),
            sim_soil(150,7.8,90, 28,"Sandy"),
            sim_soil(150,5.5,240,62,"Clay"))
df <- df[sample(nrow(df)),]

# ── Scale features ────────────────────────────────────────────
X <- scale(df[,1:3])
y <- df$soil_type

# ── Train/Test Split ──────────────────────────────────────────
idx  <- createDataPartition(y, p=0.8, list=FALSE)
X_tr <- X[idx,]; X_te <- X[-idx,]
y_tr <- y[idx];  y_te <- y[-idx]

# ── Tune k with cross-validation ─────────────────────────────
ctrl  <- trainControl(method="cv", number=5)
kgrid <- data.frame(k=1:20)
knn_cv <- train(x=X_tr, y=y_tr, method="knn",
                trControl=ctrl, tuneGrid=kgrid)
best_k <- knn_cv$bestTune$k
cat(sprintf("Best k = %d\n", best_k))

# ── Final predictions ─────────────────────────────────────────
preds <- knn(train=X_tr, test=X_te, cl=y_tr, k=best_k)
cm    <- confusionMatrix(preds, y_te)
print(cm)

# ── New sample prediction ─────────────────────────────────────
new_s <- scale(matrix(c(6.0, 175, 42), nrow=1),
               center=attr(X,"scaled:center"),
               scale=attr(X,"scaled:scale"))
pred  <- knn(train=X_tr, test=new_s, cl=y_tr, k=best_k)
cat(sprintf("\nNew sample predicted as: %s\n", pred))

1.4  Support Vector Machine (SVM)

Margin MaximisationKernel Methods
DEFINITION

SVM finds the optimal hyperplane that maximises the margin between two classes in the feature (or kernel-transformed) space. Support vectors are the training points closest to the decision boundary; all other points do not influence the boundary. The kernel trick allows SVM to separate non-linearly separable data without explicitly computing the transformed coordinates.

1.4.1  Mathematical Foundation

FORMULAE

Hard-margin optimisation:

\[\min_{\mathbf{w},b}\;\frac{1}{2}\|\mathbf{w}\|^2 \quad\text{s.t.}\quad y_i(\mathbf{w}^\top\mathbf{x}_i + b) \ge 1 \;\forall i\]

Soft-margin (with slack \(\xi_i\)):

\[\min_{\mathbf{w},b,\boldsymbol{\xi}}\;\frac{1}{2}\|\mathbf{w}\|^2 + C\sum_{i=1}^N \xi_i \quad\text{s.t.}\quad y_i(\mathbf{w}^\top\mathbf{x}_i+b)\ge 1-\xi_i,\;\xi_i\ge 0\]

Margin width: \(\dfrac{2}{\|\mathbf{w}\|}\)

Common Kernels:

\[K_{\text{RBF}}(\mathbf{x},\mathbf{z}) = \exp\!\left(-\gamma\|\mathbf{x}-\mathbf{z}\|^2\right)\] \[K_{\text{poly}}(\mathbf{x},\mathbf{z}) = (\mathbf{x}^\top\mathbf{z} + r)^d\]

Decision function: \(f(\mathbf{x}) = \text{sign}\!\left(\sum_i \alpha_i y_i K(\mathbf{x}_i,\mathbf{x}) + b\right)\)

1.4.2  How It Works

The parameter \(C\) controls the bias-variance tradeoff: large \(C\) → low bias, high variance (few misclassifications allowed); small \(C\) → wider margin tolerating some misclassifications. The RBF kernel maps data into infinite-dimensional space, enabling complex non-linear boundaries. SVM is particularly powerful when the number of features exceeds the number of samples. Multi-class extension uses One-vs-One or One-vs-Rest strategies.

1.4.3  Assumptions and Failure Modes

ASSUMES
  • A wide-margin separator exists in the feature or kernel space
BREAKS WHEN
  • Classes overlap heavily — every point becomes a support vector
  • \(N\) is large — training is roughly \(O(N^2)\) to \(O(N^3)\)
  • Raw scores are read as probabilities — they are not, without Platt scaling

1.4.4  Worked Examples

FINANCE

📈 Stock Movement Classification

Using 10 technical indicators (RSI, MACD, Bollinger Band position), an SVM classifies next-day price movement as UP or DOWN. Trained on 2 years of daily data with an RBF kernel. Be sceptical of the accuracy: next-day direction is close to a coin flip, and anything much above 55% out-of-sample usually signals look-ahead bias or a leaky train/test split rather than signal. Split time series chronologically, never at random.

RSIMACDBB_posMovement
720.420.88UP
28-0.310.12DOWN
550.050.51UP
AGRICULTURE

🛰️ Land Use Classification

Satellite multi-spectral bands (NIR, Red, Green) are used to classify land as forest, cropland, or water body. Multi-spectral land cover is unusually separable — water, vegetation and bare soil occupy distinct regions of NIR/Red space — so a well-tuned RBF SVM performs strongly here. The harder problem is spatial autocorrelation: neighbouring pixels are not independent, so random pixel splits overstate accuracy.

NIRRedGreenLand Use
0.820.140.21Forest
0.450.380.31Cropland
0.080.060.42Water
MEDICINE

🔬 Cancer Cell Classification

Gene expression profiles (around 500 genes) from a few hundred tissue biopsies classify samples as cancerous or healthy. This is the \(p \gg n\) regime where SVM is at its strongest. A linear kernel usually beats RBF here: in 500 dimensions the classes are already close to linearly separable, and the extra flexibility only buys overfitting.

Gene_AGene_BGene_CLabel
3.42-1.22.81Cancer
0.310.88-0.4Healthy

1.4.5  Code

SVM
import numpy as np
import pandas as pd
from sklearn.svm import SVC
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report, roc_auc_score
from sklearn.pipeline import Pipeline

np.random.seed(42)

# ── Simulate Land-Use Satellite Data ─────────────────────────
def make_class(n, nir, red, green, label):
    return pd.DataFrame({
        'NIR':   np.clip(np.random.normal(nir,0.06,n),0,1),
        'Red':   np.clip(np.random.normal(red,0.05,n),0,1),
        'Green': np.clip(np.random.normal(green,0.05,n),0,1),
        'label': [label]*n
    })

df = pd.concat([
    make_class(1800, 0.82, 0.14, 0.21, 'Forest'),
    make_class(2000, 0.45, 0.38, 0.31, 'Cropland'),
    make_class(1200, 0.08, 0.06, 0.42, 'Water')
]).sample(frac=1, random_state=42).reset_index(drop=True)

X = df[['NIR','Red','Green']].values
y = df['label'].values

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=42)

# ── Pipeline: Scale + SVM ─────────────────────────────────────
pipe = Pipeline([('sc', StandardScaler()), ('svm', SVC(kernel='rbf', probability=True))])

# ── Hyperparameter tuning ─────────────────────────────────────
param_grid = {'svm__C': [0.1, 1, 10, 100], 'svm__gamma': ['scale','auto',0.01,0.1]}
gs = GridSearchCV(pipe, param_grid, cv=5, scoring='accuracy', n_jobs=-1, verbose=0)
gs.fit(X_tr, y_tr)
print(f"Best params: {gs.best_params_}")

best_svm = gs.best_estimator_
y_pred = best_svm.predict(X_te)
print("\n=== SVM — Land Use Classification ===")
print(classification_report(y_te, y_pred))

# ── Decision function values (distance to hyperplane) ─────────
decision_vals = best_svm.decision_function(X_te[:5])
print("Decision function values (first 5 samples):\n", np.round(decision_vals, 3))

# ── Financial: Stock Movement ──────────────────────────────────
np.random.seed(1)
n = 500
# Simulate RSI, MACD, BB_pos for UP/DOWN
up   = np.column_stack([np.random.normal(62,10,250), np.random.normal(0.3,0.2,250), np.random.normal(0.68,0.15,250)])
down = np.column_stack([np.random.normal(38,10,250), np.random.normal(-0.3,0.2,250), np.random.normal(0.32,0.15,250)])
Xs   = np.vstack([up, down])
ys   = np.array([1]*250 + [0]*250)
Xs_tr, Xs_te, ys_tr, ys_te = train_test_split(Xs, ys, test_size=0.2, random_state=1)
pipe2 = Pipeline([('sc', StandardScaler()), ('svm', SVC(C=1, kernel='rbf', probability=True))])
pipe2.fit(Xs_tr, ys_tr)
auc = roc_auc_score(ys_te, pipe2.predict_proba(Xs_te)[:,1])
print(f"\nStock Direction AUC-ROC: {auc:.4f}")
library(e1071); library(caret); set.seed(42)

# ── Simulate Land-Use Data ────────────────────────────────────
make_cls <- function(n, nir, red, green, lbl) {
  data.frame(NIR=rnorm(n,nir,.06), Red=rnorm(n,red,.05),
             Green=rnorm(n,green,.05), label=factor(rep(lbl,n)))
}
df <- rbind(make_cls(1800,.82,.14,.21,"Forest"),
            make_cls(2000,.45,.38,.31,"Cropland"),
            make_cls(1200,.08,.06,.42,"Water"))
df <- df[sample(nrow(df)),]

idx  <- createDataPartition(df$label, p=0.8, list=FALSE)
tr   <- df[idx,]; te <- df[-idx,]

# ── SVM with RBF kernel ───────────────────────────────────────
svm_model <- svm(label ~ ., data=tr, kernel="radial",
                 cost=10, gamma=0.5, probability=TRUE)

preds <- predict(svm_model, te)
cm    <- confusionMatrix(preds, te$label)
cat("=== SVM — Land Use Classification ===\n"); print(cm)

# ── Tune via caret ────────────────────────────────────────────
ctrl <- trainControl(method="cv", number=3)
grid <- expand.grid(C=c(1,10), sigma=c(0.01,0.1))
svm_cv <- train(label~., data=tr, method="svmRadial",
                trControl=ctrl, tuneGrid=grid, preProc=c("center","scale"))
cat(sprintf("\nBest C=%.0f, sigma=%.3f, Accuracy=%.4f\n",
            svm_cv$bestTune$C, svm_cv$bestTune$sigma,
            max(svm_cv$results$Accuracy)))

1.5  Decision Tree (Classification)

Tree-BasedInterpretable
DEFINITION

A Decision Tree recursively partitions the feature space using axis-aligned splits to create a tree where each internal node tests a feature, each branch represents an outcome, and each leaf holds a class prediction. The tree is built greedily by selecting the split that maximises information gain (or minimises Gini impurity).

1.5.1  Mathematical Foundation

FORMULAE

Gini Impurity:

\[G(D) = 1 - \sum_{k=1}^{K} p_k^2\]

Entropy:

\[H(D) = -\sum_{k=1}^{K} p_k \log_2 p_k\]

Information Gain for split \(s\) on node \(D\):

\[IG(D, s) = H(D) - \sum_{j \in \{L,R\}} \frac{|D_j|}{|D|}\, H(D_j)\]

Best split: \(s^* = \arg\max_s \; IG(D, s)\)

Pruning (cost-complexity):

\[R_\alpha(T) = R(T) + \alpha\,|T_{\text{leaves}}|\]

1.5.2  How It Works

CART (Classification And Regression Trees) uses binary splits with Gini impurity. ID3 uses entropy with information gain and splits multiway (one branch per category), so the sum above runs over all children, not just \(\{L,R\}\). C4.5 replaces information gain with the gain ratio, \(IG(D,s)/H(s)\), precisely to correct IG's bias toward high-cardinality splits — a feature with a unique value per row scores perfect information gain and is useless. The tree grows until a stopping criterion (max depth, min samples per leaf) is met. Post-pruning removes branches that provide little power, controlled by \(\alpha\). Decision trees are inherently interpretable — rules can be read directly from root to leaf. They are the base learner in Random Forests and Gradient Boosted Trees. The same tree also does regression: swap Gini or entropy for variance reduction as the split criterion, and have each leaf predict the mean of its training targets rather than a majority class. Only the impurity measure and the leaf rule change; the recursive-partitioning algorithm is identical (DecisionTreeRegressor in scikit-learn).

1.5.3  Assumptions and Failure Modes

ASSUMES
  • The boundary is well approximated by axis-aligned splits
BREAKS WHEN
  • Unpruned — a single tree memorises the training set
  • The boundary is diagonal — the tree approximates it as a staircase
  • Impurity importances inflate high-cardinality features

1.5.4  Worked Examples

FINANCE

🏛️ Loan Approval Rules

A decision tree generates human-readable rules: IF credit_score > 650 AND income > 40k THEN Approve. Bank compliance officers can audit every decision path directly from the tree structure.

CreditScoreIncome $kEmployedDecision
72085YesApprove
51028NoDeny
64055YesApprove
AGRICULTURE

🌦️ Crop Recommendation

A decision tree recommends the best crop based on soil nutrients and climate: IF N>150 AND pH∈[6,7] AND rainfall>800mm THEN wheat; ELSE maize. Interpretable by non-technical farmers.

N mg/kgpHRain mmCrop
1806.5920Wheat
957.2650Maize
2105.81100Rice
MEDICINE

🏥 Emergency Triage

ER triage tree classifies patient urgency (critical, urgent, non-urgent) using heart rate, blood pressure, and SpO₂. Clinicians value the transparent rule set for audit and training.

HR bpmBP sysSpO₂%Triage
1428288Critical
9813097Urgent
7211899Non-urgent

1.5.5  Code

Decision Tree
import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split
from sklearn.metrics import classification_report
from sklearn.preprocessing import LabelEncoder

np.random.seed(42)

# ── Simulate Crop Recommendation Dataset ─────────────────────
def make_crop(n, N_mu, pH_mu, rain_mu, label):
    return pd.DataFrame({
        'N':      np.random.normal(N_mu,  20,  n),
        'pH':     np.random.normal(pH_mu, 0.3, n),
        'rain':   np.random.normal(rain_mu,80, n),
        'crop':   [label]*n
    })

df = pd.concat([
    make_crop(300, 180, 6.4, 920,  'Wheat'),
    make_crop(280, 95,  7.0, 620,  'Maize'),
    make_crop(250, 220, 5.7, 1150, 'Rice')
]).sample(frac=1, random_state=42).reset_index(drop=True)

X = df[['N','pH','rain']].values
le = LabelEncoder()
y = le.fit_transform(df['crop'])

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)

# ── Train Decision Tree (Gini) ────────────────────────────────
dt = DecisionTreeClassifier(criterion='gini', max_depth=4,
                             min_samples_leaf=10, random_state=42)
dt.fit(X_tr, y_tr)

y_pred = dt.predict(X_te)
print("=== Decision Tree — Crop Recommendation ===")
print(classification_report(y_te, y_pred, target_names=le.classes_))
print(f"Tree depth: {dt.get_depth()}, Leaves: {dt.get_n_leaves()}")

# ── Print human-readable rules ────────────────────────────────
rules = export_text(dt, feature_names=['N','pH','rain'])
print("\n--- Decision Rules ---")
print(rules)

# ── Feature Importance ────────────────────────────────────────
for name, imp in zip(['N','pH','rain'], dt.feature_importances_):
    print(f"  {name:6s}: {imp:.4f}")

# ── Financial: Loan Approval ──────────────────────────────────
np.random.seed(5)
approved = np.column_stack([np.random.normal(690,55,400), np.random.normal(72,18,400), np.ones(400)])
denied   = np.column_stack([np.random.normal(515,60,200), np.random.normal(34,12,200), np.zeros(200)])
Xf = np.vstack([approved[:,:2], denied[:,:2]])
yf = np.array([1]*400 + [0]*200)
Xf_tr, Xf_te, yf_tr, yf_te = train_test_split(Xf, yf, test_size=0.2, random_state=5)
dt_fin = DecisionTreeClassifier(max_depth=3, random_state=5).fit(Xf_tr, yf_tr)
print(f"\nLoan Approval Tree Accuracy: {dt_fin.score(Xf_te, yf_te):.4f}")
print(export_text(dt_fin, feature_names=['CreditScore','Income_k']))
library(rpart); library(rpart.plot); library(caret); set.seed(42)

# ── Simulate Crop Dataset ─────────────────────────────────────
make_crop <- function(n, N_mu, pH_mu, rain_mu, label)
  data.frame(N=rnorm(n,N_mu,20), pH=rnorm(n,pH_mu,.3),
             rain=rnorm(n,rain_mu,80), crop=factor(rep(label,n)))

df <- rbind(make_crop(300,180,6.4,920,"Wheat"),
            make_crop(280,95, 7.0,620,"Maize"),
            make_crop(250,220,5.7,1150,"Rice"))
df <- df[sample(nrow(df)),]

idx <- createDataPartition(df$crop, p=0.8, list=FALSE)
tr  <- df[idx,]; te <- df[-idx,]

# ── Train CART tree ───────────────────────────────────────────
tree <- rpart(crop ~ ., data=tr, method="class",
              control=rpart.control(minsplit=20, maxdepth=4, cp=0.01))

# ── Print rules ───────────────────────────────────────────────
cat("=== Decision Rules ===\n")
print(tree)

# ── Evaluate ──────────────────────────────────────────────────
preds <- predict(tree, te, type="class")
cm    <- confusionMatrix(preds, te$crop)
print(cm)

# ── Feature importance ────────────────────────────────────────
cat("\nVariable Importance:\n")
print(tree$variable.importance)

# ── Cost-complexity pruning ───────────────────────────────────
printcp(tree)
best_cp <- tree$cptable[which.min(tree$cptable[,"xerror"]),"CP"]
pruned  <- prune(tree, cp=best_cp)
cat(sprintf("\nPruned tree nodes: %d  (best cp=%.5f)\n",
            nrow(pruned$frame), best_cp))

At a Glance

The same information as the assumption blocks above, side by side — this is the comparison that decides which method to reach for.

AlgorithmAssumesBreaks when
1.1 Naive Bayes ClassifierFeatures conditionally independent given the classFeatures are strongly correlated — posteriors become badly over-confident
1.2 Logistic RegressionLog-odds linear in the featuresClasses are perfectly separable — coefficients diverge without regularisation
1.3 K-Nearest Neighbors (KNN)Nearby points share labelsFeatures are left unscaled — the largest-range feature dominates the distance
1.4 Support Vector Machine (SVM)A wide-margin separator exists in the feature or kernel spaceClasses overlap heavily — every point becomes a support vector
1.5 Decision Tree (Classification)The boundary is well approximated by axis-aligned splitsUnpruned — a single tree memorises the training set