Skip to the content

Algorithms on This Page

2.6 Isolation Forest
A boundary case worth naming: anomaly detection spans all three supervision regimes. It is unsupervised when no labels exist and rarity itself defines the anomaly — the setting of Isolation Forest below. It is semi-supervised when a clean sample of normal data is available to model, and supervised when confirmed anomalies are labelled, at which point it becomes imbalanced binary classification. It is placed here because the unsupervised framing needs the fewest assumptions, not because the other two are wrong.

2.6  Isolation Forest

Anomaly DetectionTree Ensemble
DEFINITION

Isolation Forest detects anomalies by isolating observations: it randomly selects a feature and a random split value, recursively partitioning the data. Anomalies are rare and different, so they are isolated in fewer steps (shorter average path length in isolation trees) compared to normal points which require many splits to isolate.

2.6.1  Mathematical Foundation

FORMULAE

Average path length normalisation:

\[c(n) = 2H(n-1) - \frac{2(n-1)}{n}, \quad H(i) = \ln(i) + \gamma_{\text{Euler}}\]

Anomaly Score (higher → more anomalous):

\[s(\mathbf{x}, n) = 2^{-\frac{E[h(\mathbf{x})]}{c(n)}}\]

where \(h(\mathbf{x})\) is the path length of \(\mathbf{x}\) in an isolation tree

Reading the score (Liu, Ting & Zhou, 2008):

\(s \to 1\) — path length far below average: very likely an anomaly
\(s \ll 0.5\) — path length far above average: safely normal
\(s \approx 0.5\) for every point — a statement about the sample, not the point: the data contains no distinct anomaly

Decision boundary — two equivalent ways to threshold:

By score: flag when \(s > s_0\) (commonly \(s_0 \approx 0.6\)). By quantile: flag the top \(k\%\) of scores. scikit-learn takes the second route — its contamination parameter is \(k\), and it sets the internal offset_ so that exactly that fraction is flagged. Note that score_samples() returns \(-s\), so -iso.score_samples(X) recovers \(s\) as defined above.

2.6.2  How It Works

Isolation Forest builds \(T\) isolation trees (each on a subsample), then averages path lengths across trees. The key insight is that anomalies are few and far from normal regions, so random partitioning isolates them quickly. The algorithm is linear in time complexity \(O(Tn)\) and memory, making it highly scalable. The contamination parameter sets the expected proportion of anomalies. Extended Isolation Forest addresses the axis-aligned partition bias of the original algorithm.

2.6.3  Assumptions and Failure Modes

ASSUMES
  • Anomalies are few, and different enough to isolate in few splits
BREAKS WHEN
  • Anomalies form their own dense cluster — they stop being easy to isolate
  • contamination is set from the true rate — that leaks the answer
  • Anomalies are defined by feature combinations, not axis-aligned cuts

2.6.4  Worked Examples

FINANCE

💳 Real-Time Fraud Detection

Processing 100,000 daily credit card transactions, Isolation Forest flags the top 1% (1,000 transactions) as anomalous. The precision/recall trade-off moves entirely with the assumed contamination rate, which in deployment is unknown. The code below sweeps that assumption so the trade-off is visible rather than hidden inside one flattering number.

ScoreAmount $LocationLabel
0.913,420ForeignFraud
0.871,800DomesticReview
0.3145LocalNormal
AGRICULTURE

🌡️ Sensor Malfunction Detection

IoT weather station network with 200 sensors. Isolation Forest detects 8 sensors with anomalous temperature/humidity readings (due to hardware faults or vandalism) before they corrupt regional weather models.

SensorTemp°CHumidity%Status
S_042-45102Anomaly
S_1078812Anomaly
S_0151868Normal
MEDICINE

⚠️ Adverse Drug Event Detection

Monitoring 50,000 post-marketing drug safety reports. Isolation Forest identifies rare adverse event patterns (e.g., sudden liver enzyme spike + rash within 48h of drug start) for immediate pharmacovigilance review.

ALT U/LRashDays post-RxFlag
420Yes2Anomaly
28No14Normal

2.6.5  Code

Isolation Forest
import numpy as np
import pandas as pd
from sklearn.ensemble import IsolationForest
from sklearn.metrics import precision_score, recall_score, f1_score, roc_auc_score
from sklearn.preprocessing import StandardScaler

np.random.seed(42)

# ── Simulate Credit Card Transactions ─────────────────────────
n_normal, n_fraud = 9800, 200

normal = pd.DataFrame({
    'amount':   np.abs(np.random.normal(85, 60, n_normal)),
    'freq_day': np.random.poisson(3, n_normal).astype(float),
    'loc_dev':  np.abs(np.random.normal(5, 10, n_normal)),
    'hour':     np.random.randint(7, 23, n_normal).astype(float),
    'label':    0
})
fraud = pd.DataFrame({
    'amount':   np.abs(np.random.normal(1800, 800, n_fraud)),
    'freq_day': np.random.poisson(18, n_fraud).astype(float),
    'loc_dev':  np.abs(np.random.normal(900, 300, n_fraud)),
    'hour':     np.random.randint(0, 5, n_fraud).astype(float),
    'label':    1
})
df = pd.concat([normal, fraud]).sample(frac=1, random_state=42).reset_index(drop=True)

X = df[['amount','freq_day','loc_dev','hour']].values
y = df['label'].values

sc = StandardScaler()
X_sc = sc.fit_transform(X)

# ── Isolation Forest ──────────────────────────────────────────
# `contamination` is the ASSUMED anomaly rate and sets the flagging threshold.
# In deployment it is unknown — setting it to the true rate (here 0.02) would
# hand the model the answer it is supposed to estimate. Use 'auto', which
# applies the offset from the original paper, and inspect the operating
# point separately below.
iso = IsolationForest(n_estimators=200, contamination='auto',
                       max_samples='auto', random_state=42)
iso.fit(X_sc)

scores      = -iso.score_samples(X_sc)  # higher = more anomalous
predictions = iso.predict(X_sc)         # -1 = anomaly, +1 = normal
y_pred      = (predictions == -1).astype(int)

print("=== Isolation Forest — Credit Card Fraud ===")
print("Operating point: contamination='auto' — the true 2% rate is NOT supplied.")
print(f"Anomalies detected : {y_pred.sum()} / {len(y_pred)}")
print(f"Precision          : {precision_score(y, y_pred):.4f}")
print(f"Recall             : {recall_score(y, y_pred):.4f}")
print(f"F1-Score           : {f1_score(y, y_pred):.4f}")
print(f"AUC-ROC            : {roc_auc_score(y, scores):.4f}")

# ── Anomaly score distribution ────────────────────────────────
normal_scores = scores[y == 0]
fraud_scores  = scores[y == 1]
print(f"\nAnomaly Score — Normal: mean={normal_scores.mean():.4f}, std={normal_scores.std():.4f}")
print(f"Anomaly Score — Fraud : mean={fraud_scores.mean():.4f}, std={fraud_scores.std():.4f}")

# ── Effect of contamination parameter ────────────────────────
# This is the whole story: AUC-ROC above is threshold-free and stays perfect,
# while precision/recall swing wildly with an assumption you cannot verify
# without labels. Quote the sweep, never a single flattering F1.
print("\nEffect of contamination parameter:")
for cont in [0.005, 0.01, 0.02, 0.05, 0.10]:
    m = IsolationForest(n_estimators=100, contamination=cont, random_state=42)
    p = (m.fit_predict(X_sc) == -1).astype(int)
    print(f"  contamination={cont:.3f}: flagged={p.sum():5d}, F1={f1_score(y,p):.4f}")
library(isotree); set.seed(42)

# ── Simulate Transaction Data ─────────────────────────────────
n_norm <- 9800; n_fraud <- 200
normal <- data.frame(amount=abs(rnorm(n_norm,85,60)),
                     freq_day=rpois(n_norm,3)+0.0,
                     loc_dev=abs(rnorm(n_norm,5,10)),
                     hour=runif(n_norm,7,23),
                     label=0)
fraud  <- data.frame(amount=abs(rnorm(n_fraud,1800,800)),
                     freq_day=rpois(n_fraud,18)+0.0,
                     loc_dev=abs(rnorm(n_fraud,900,300)),
                     hour=runif(n_fraud,0,5),
                     label=1)
df <- rbind(normal, fraud)[sample(n_norm+n_fraud),]

X <- scale(df[,1:4])
y <- df$label

# ── Isolation Forest ──────────────────────────────────────────
iso_model <- isolation.forest(as.data.frame(X),
                               ntrees=200, sample_size=256,
                               nthreads=1)
scores    <- predict(iso_model, as.data.frame(X))

threshold <- quantile(scores, 0.98)  # top 2% as anomaly
y_pred    <- as.integer(scores >= threshold)

# ── Metrics ───────────────────────────────────────────────────
tp <- sum(y_pred==1 & y==1); fp <- sum(y_pred==1 & y==0)
fn <- sum(y_pred==0 & y==1)
prec <- tp/(tp+fp); rec <- tp/(tp+fn); f1 <- 2*prec*rec/(prec+rec)
cat(sprintf("Precision=%.4f  Recall=%.4f  F1=%.4f\n", prec, rec, f1))

# ── AUC ───────────────────────────────────────────────────────
library(pROC)
roc_obj <- roc(y, scores, quiet=TRUE)
cat(sprintf("AUC-ROC = %.4f\n", auc(roc_obj)))

cat(sprintf("Normal score:  mean=%.4f\nFraud score:   mean=%.4f\n",
            mean(scores[y==0]), mean(scores[y==1])))

At a Glance

The same information as the assumption blocks above, side by side — this is the comparison that decides which method to reach for.

AlgorithmAssumesBreaks when
2.6 Isolation ForestAnomalies are few, and different enough to isolate in few splitsAnomalies form their own dense cluster — they stop being easy to isolate