Isolation Forest detects anomalies by isolating observations: it randomly selects a feature and a random split value, recursively partitioning the data. Anomalies are rare and different, so they are isolated in fewer steps (shorter average path length in isolation trees) compared to normal points which require many splits to isolate.
Average path length normalisation:
\[c(n) = 2H(n-1) - \frac{2(n-1)}{n}, \quad H(i) = \ln(i) + \gamma_{\text{Euler}}\]Anomaly Score (higher → more anomalous):
\[s(\mathbf{x}, n) = 2^{-\frac{E[h(\mathbf{x})]}{c(n)}}\]where \(h(\mathbf{x})\) is the path length of \(\mathbf{x}\) in an isolation tree
Reading the score (Liu, Ting & Zhou, 2008):
\(s \to 1\) — path length far below average: very likely an anomaly
\(s \ll 0.5\) — path length far above average: safely normal
\(s \approx 0.5\) for every point — a statement about the sample, not the point: the data contains no distinct anomaly
Decision boundary — two equivalent ways to threshold:
By score: flag when \(s > s_0\) (commonly \(s_0 \approx 0.6\)). By quantile: flag the top \(k\%\) of scores. scikit-learn takes the second route — its contamination parameter is \(k\), and it sets the internal offset_ so that exactly that fraction is flagged. Note that score_samples() returns \(-s\), so -iso.score_samples(X) recovers \(s\) as defined above.
Isolation Forest builds \(T\) isolation trees (each on a subsample), then averages path lengths across trees. The key insight is that anomalies are few and far from normal regions, so random partitioning isolates them quickly. The algorithm is linear in time complexity \(O(Tn)\) and memory, making it highly scalable. The contamination parameter sets the expected proportion of anomalies. Extended Isolation Forest addresses the axis-aligned partition bias of the original algorithm.
contamination is set from the true rate — that leaks the answerProcessing 100,000 daily credit card transactions, Isolation Forest flags the top 1% (1,000 transactions) as anomalous. The precision/recall trade-off moves entirely with the assumed contamination rate, which in deployment is unknown. The code below sweeps that assumption so the trade-off is visible rather than hidden inside one flattering number.
| Score | Amount $ | Location | Label |
|---|---|---|---|
| 0.91 | 3,420 | Foreign | Fraud |
| 0.87 | 1,800 | Domestic | Review |
| 0.31 | 45 | Local | Normal |
IoT weather station network with 200 sensors. Isolation Forest detects 8 sensors with anomalous temperature/humidity readings (due to hardware faults or vandalism) before they corrupt regional weather models.
| Sensor | Temp°C | Humidity% | Status |
|---|---|---|---|
| S_042 | -45 | 102 | Anomaly |
| S_107 | 88 | 12 | Anomaly |
| S_015 | 18 | 68 | Normal |
Monitoring 50,000 post-marketing drug safety reports. Isolation Forest identifies rare adverse event patterns (e.g., sudden liver enzyme spike + rash within 48h of drug start) for immediate pharmacovigilance review.
| ALT U/L | Rash | Days post-Rx | Flag |
|---|---|---|---|
| 420 | Yes | 2 | Anomaly |
| 28 | No | 14 | Normal |
import numpy as np
import pandas as pd
from sklearn.ensemble import IsolationForest
from sklearn.metrics import precision_score, recall_score, f1_score, roc_auc_score
from sklearn.preprocessing import StandardScaler
np.random.seed(42)
# ── Simulate Credit Card Transactions ─────────────────────────
n_normal, n_fraud = 9800, 200
normal = pd.DataFrame({
'amount': np.abs(np.random.normal(85, 60, n_normal)),
'freq_day': np.random.poisson(3, n_normal).astype(float),
'loc_dev': np.abs(np.random.normal(5, 10, n_normal)),
'hour': np.random.randint(7, 23, n_normal).astype(float),
'label': 0
})
fraud = pd.DataFrame({
'amount': np.abs(np.random.normal(1800, 800, n_fraud)),
'freq_day': np.random.poisson(18, n_fraud).astype(float),
'loc_dev': np.abs(np.random.normal(900, 300, n_fraud)),
'hour': np.random.randint(0, 5, n_fraud).astype(float),
'label': 1
})
df = pd.concat([normal, fraud]).sample(frac=1, random_state=42).reset_index(drop=True)
X = df[['amount','freq_day','loc_dev','hour']].values
y = df['label'].values
sc = StandardScaler()
X_sc = sc.fit_transform(X)
# ── Isolation Forest ──────────────────────────────────────────
# `contamination` is the ASSUMED anomaly rate and sets the flagging threshold.
# In deployment it is unknown — setting it to the true rate (here 0.02) would
# hand the model the answer it is supposed to estimate. Use 'auto', which
# applies the offset from the original paper, and inspect the operating
# point separately below.
iso = IsolationForest(n_estimators=200, contamination='auto',
max_samples='auto', random_state=42)
iso.fit(X_sc)
scores = -iso.score_samples(X_sc) # higher = more anomalous
predictions = iso.predict(X_sc) # -1 = anomaly, +1 = normal
y_pred = (predictions == -1).astype(int)
print("=== Isolation Forest — Credit Card Fraud ===")
print("Operating point: contamination='auto' — the true 2% rate is NOT supplied.")
print(f"Anomalies detected : {y_pred.sum()} / {len(y_pred)}")
print(f"Precision : {precision_score(y, y_pred):.4f}")
print(f"Recall : {recall_score(y, y_pred):.4f}")
print(f"F1-Score : {f1_score(y, y_pred):.4f}")
print(f"AUC-ROC : {roc_auc_score(y, scores):.4f}")
# ── Anomaly score distribution ────────────────────────────────
normal_scores = scores[y == 0]
fraud_scores = scores[y == 1]
print(f"\nAnomaly Score — Normal: mean={normal_scores.mean():.4f}, std={normal_scores.std():.4f}")
print(f"Anomaly Score — Fraud : mean={fraud_scores.mean():.4f}, std={fraud_scores.std():.4f}")
# ── Effect of contamination parameter ────────────────────────
# This is the whole story: AUC-ROC above is threshold-free and stays perfect,
# while precision/recall swing wildly with an assumption you cannot verify
# without labels. Quote the sweep, never a single flattering F1.
print("\nEffect of contamination parameter:")
for cont in [0.005, 0.01, 0.02, 0.05, 0.10]:
m = IsolationForest(n_estimators=100, contamination=cont, random_state=42)
p = (m.fit_predict(X_sc) == -1).astype(int)
print(f" contamination={cont:.3f}: flagged={p.sum():5d}, F1={f1_score(y,p):.4f}")
library(isotree); set.seed(42)
# ── Simulate Transaction Data ─────────────────────────────────
n_norm <- 9800; n_fraud <- 200
normal <- data.frame(amount=abs(rnorm(n_norm,85,60)),
freq_day=rpois(n_norm,3)+0.0,
loc_dev=abs(rnorm(n_norm,5,10)),
hour=runif(n_norm,7,23),
label=0)
fraud <- data.frame(amount=abs(rnorm(n_fraud,1800,800)),
freq_day=rpois(n_fraud,18)+0.0,
loc_dev=abs(rnorm(n_fraud,900,300)),
hour=runif(n_fraud,0,5),
label=1)
df <- rbind(normal, fraud)[sample(n_norm+n_fraud),]
X <- scale(df[,1:4])
y <- df$label
# ── Isolation Forest ──────────────────────────────────────────
iso_model <- isolation.forest(as.data.frame(X),
ntrees=200, sample_size=256,
nthreads=1)
scores <- predict(iso_model, as.data.frame(X))
threshold <- quantile(scores, 0.98) # top 2% as anomaly
y_pred <- as.integer(scores >= threshold)
# ── Metrics ───────────────────────────────────────────────────
tp <- sum(y_pred==1 & y==1); fp <- sum(y_pred==1 & y==0)
fn <- sum(y_pred==0 & y==1)
prec <- tp/(tp+fp); rec <- tp/(tp+fn); f1 <- 2*prec*rec/(prec+rec)
cat(sprintf("Precision=%.4f Recall=%.4f F1=%.4f\n", prec, rec, f1))
# ── AUC ───────────────────────────────────────────────────────
library(pROC)
roc_obj <- roc(y, scores, quiet=TRUE)
cat(sprintf("AUC-ROC = %.4f\n", auc(roc_obj)))
cat(sprintf("Normal score: mean=%.4f\nFraud score: mean=%.4f\n",
mean(scores[y==0]), mean(scores[y==1])))
The same information as the assumption blocks above, side by side — this is the comparison that decides which method to reach for.
| Algorithm | Assumes | Breaks when |
|---|---|---|
| 2.6 Isolation Forest | Anomalies are few, and different enough to isolate in few splits | Anomalies form their own dense cluster — they stop being easy to isolate |