Self-Training (Self-Learning) is the simplest semi-supervised method: a classifier trained on labelled data makes predictions on unlabelled data, and the highest-confidence predictions become pseudo-labels added to the training set. The model is retrained iteratively, gradually incorporating more of the unlabelled pool.
Algorithm:
Confidence threshold: \(\tau \in [0.8, 0.99]\) — higher reduces noise in pseudo-labels
Risk: Confirmation bias — errors compound if initial model is poor on \(\mathcal{L}\)
Self-training rests on confidence calibration — the assumption that a high predicted probability really does track a high chance of being right. (This is distinct from the low-density separation assumption, a geometric claim that the decision boundary passes through sparse regions, which is what justifies transductive SVMs.) Calibration is why confirmation bias is the characteristic failure mode: an over-confident model pseudo-labels its own errors and then trains on them. The choice of threshold \(\tau\) is critical: too low introduces label noise, too high slows progress. Class-balanced selection at each iteration prevents class imbalance propagation. Self-training is particularly effective when the labelled set is small but representative, and the unlabelled set is large and similar in distribution.
A new bank has only 200 labelled loan accounts (default/no default) but 5,000 unlabelled historical accounts. Self-training expands the effective training set with the pseudo-labels the model is most confident about. The gain is real when the unlabelled pool follows the same distribution as the labelled one — and can be negative when it does not, because early errors are trained on and compound.
| Iteration | Labelled | AUC |
|---|---|---|
| 0 (initial) | 200 | 0.71 |
| 3 | 820 | 0.78 |
| 10 | 1,800 | 0.83 |
Only 50 satellite image pixels have been manually labelled as crop type. Self-training with a Random Forest iteratively labels the pool, starting with the pixels most like the 50 labelled by hand. The confidence threshold \(\tau\) is what keeps an early mistake from propagating through the remaining 9,950.
| Method | Labelled used | Accuracy |
|---|---|---|
| Supervised only | 50 | 72% |
| Self-training | 50+9,500 | 88% |
100 patient records manually annotated for sepsis vs non-sepsis. Self-training on 10,000 unannotated records is attractive because clinical annotation, not data, is the bottleneck. But a model that is over-confident on rare sepsis cases will cheerfully pseudo-label its own misses — so calibration matters more here than headline accuracy.
| Method | F1-Score |
|---|---|
| Supervised (100) | 0.61 |
| Self-training | 0.78 |
| Full supervision | 0.86 |
import numpy as np
import pandas as pd
from sklearn.semi_supervised import SelfTrainingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report, roc_auc_score
from sklearn.model_selection import train_test_split
np.random.seed(42)
N = 5000
# ── Simulate Loan Default Dataset ─────────────────────────────
X_all = np.column_stack([
np.random.normal(0, 1, N), # credit score (std)
np.random.normal(0, 1, N), # income (std)
np.random.normal(0, 1, N) # debt ratio (std)
])
true_y = (X_all[:,0] - X_all[:,2] + 0.5*X_all[:,1] +
np.random.normal(0, 0.5, N) < 0).astype(int)
# ── Only 200 labelled; rest unlabelled (-1) ───────────────────
X_tr_all, X_te, y_tr_all, y_te = train_test_split(X_all, true_y, test_size=0.2, random_state=42)
n_labelled = 200
y_semi = y_tr_all.copy().astype(float)
unlabelled_idx = np.random.choice(len(y_semi), len(y_semi)-n_labelled, replace=False)
y_semi[unlabelled_idx] = -1 # -1 means unlabelled in sklearn
scaler = StandardScaler()
X_tr_sc = scaler.fit_transform(X_tr_all)
X_te_sc = scaler.transform(X_te)
# ── Supervised baseline (labelled only) ───────────────────────
labelled_mask = y_semi != -1
lr_sup = LogisticRegression(max_iter=1000).fit(X_tr_sc[labelled_mask], y_semi[labelled_mask])
auc_sup = roc_auc_score(y_te, lr_sup.predict_proba(X_te_sc)[:,1])
# ── Self-Training ─────────────────────────────────────────────
base = LogisticRegression(max_iter=1000, random_state=42)
# NOTE: the first argument was renamed `base_estimator` -> `estimator` in
# scikit-learn 1.6. Pass it positionally and the code works on both.
st = SelfTrainingClassifier(base, threshold=0.85,
criterion='threshold', max_iter=20, verbose=False)
st.fit(X_tr_sc, y_semi)
auc_st = roc_auc_score(y_te, st.predict_proba(X_te_sc)[:,1])
print("=== Self-Training — Loan Default ===")
print(f"Supervised (200 labels) AUC : {auc_sup:.4f}")
print(f"Self-Training AUC : {auc_st:.4f}")
print(f"Iterations used : {st.n_iter_}")
print(f"Pseudo-labelled samples : {(st.transduction_ != -1).sum() - n_labelled}")
print("\nClassification Report (Self-Training):")
print(classification_report(y_te, st.predict(X_te_sc), target_names=['No Default','Default']))
# ── Effect of threshold ───────────────────────────────────────
print("\nEffect of confidence threshold:")
for tau in [0.70, 0.80, 0.85, 0.90, 0.95]:
m = SelfTrainingClassifier(LogisticRegression(max_iter=1000), threshold=tau)
m.fit(X_tr_sc, y_semi)
a = roc_auc_score(y_te, m.predict_proba(X_te_sc)[:,1])
print(f" τ={tau:.2f} AUC={a:.4f} iters={m.n_iter_}")
library(RSSL); set.seed(42)
# ── Simulate Dataset ──────────────────────────────────────────
N <- 2000
X1 <- rnorm(N); X2 <- rnorm(N); X3 <- rnorm(N)
y <- factor(as.integer(X1 - X3 + 0.5*X2 + rnorm(N,0,.5) < 0))
df <- data.frame(X1,X2,X3,y)
# ── Split: 80% train, 20% test ────────────────────────────────
idx <- sample(1:N, 0.8*N)
tr <- df[idx,]; te <- df[-idx,]
# ── Simulate only 200 labelled ───────────────────────────────
labelled_idx <- sample(1:nrow(tr), 200)
unlabelled_idx <- setdiff(1:nrow(tr), labelled_idx)
X_l <- as.matrix(tr[labelled_idx, 1:3])
y_l <- tr$y[labelled_idx]
X_u <- as.matrix(tr[unlabelled_idx, 1:3])
X_te <- as.matrix(te[,1:3])
y_te <- te$y
# ── Supervised baseline ───────────────────────────────────────
sup_model <- glm(y ~ X1+X2+X3, data=tr[labelled_idx,], family=binomial())
p_sup <- predict(sup_model, te, type="response")
auc_sup <- mean(outer(p_sup[y_te==1], p_sup[y_te==0], ">"))
cat(sprintf("Supervised AUC: %.4f\n", auc_sup))
# ── Self-Training with SVM ────────────────────────────────────
st_model <- SelfLearning(X_l, y_l, X_u,
method=LeastSquaresClassifier())
preds_st <- predict(st_model, X_te)
acc_st <- mean(preds_st == y_te)
cat(sprintf("Self-Training Accuracy: %.4f\n", acc_st))
Co-Training trains two (or more) classifiers on different views (disjoint feature subsets) of the data. Each classifier labels the most confidently-predicted unlabelled examples and passes them to the other classifier for training. The classifiers "teach" each other, leveraging the multi-view redundancy in the data.
Two-view assumption: Features can be split into \(\mathbf{x} = (\mathbf{x}^{(1)}, \mathbf{x}^{(2)})\) where each view is sufficient for classification
Algorithm:
Key assumption: Views are conditionally independent given the class label
Theoretical guarantee (Blum & Mitchell, 1998): if the two views are conditionally independent given the label, and a weak learner on one view alone achieves error slightly better than chance, then unlabelled data can be used to boost the pair to arbitrarily low error. The result is PAC-style learnability, not a per-iteration error bound — in practice the conditional-independence premise rarely holds exactly, and performance degrades gracefully as it is violated.
Co-Training requires a natural feature split into two sufficient views. In text classification, view 1 might be content words and view 2 hyperlinks. When natural views don't exist, random feature splits or different algorithms trained on the same features can substitute (Democratic Co-Training). Adding small amounts of confident pseudo-labels per iteration prevents rapid error propagation. Co-Training can outperform self-training when the two views are truly complementary.
View 1: NLP features from news articles; View 2: financial ratios (P/E, EPS growth). Co-training classifies stocks as BUY/HOLD/SELL. The two views are genuinely complementary — news sentiment and fundamentals fail in different ways and on different days — which is the condition co-training needs in order to help at all.
| Method | Accuracy |
|---|---|
| NLP view only | 65% |
| Fundamental view | 61% |
| Co-Training | 74% |
View 1: optical satellite bands (RGB, NIR); View 2: SAR radar backscatter. Co-training leverages weather-independent SAR and detail-rich optical data jointly, from only 80 labelled fields. SAR sees through cloud while optical carries texture, so the two views' errors are close to independent — precisely the premise Blum & Mitchell's guarantee rests on.
| Method | Labels | Accuracy |
|---|---|---|
| Optical only | 80 | 79% |
| Co-Training | 80+4,920 | 92% |
View 1: chest X-ray features (CNN embeddings); View 2: clinical note NLP features. Co-training classifies pneumonia from 50 labelled cases and 2,000 unlabelled records — a setting that matters most where radiologists are scarce. Image features and note features are close to conditionally independent given the diagnosis, which is the assumption the method needs.
| Method | Sensitivity |
|---|---|
| X-ray only (50) | 68% |
| Co-Training | 85% |
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
np.random.seed(42)
N = 3000
# ── Simulate Two-View Data (Financial: NLP + Fundamentals) ────
# View 1: NLP-like features (3 dims)
# View 2: Fundamental ratios (3 dims)
# True signal split across both views
signal = np.random.normal(0, 1, N)
view1 = np.column_stack([signal + np.random.normal(0, 0.8, N),
np.random.normal(0, 1, N),
np.random.normal(0, 1, N)])
view2 = np.column_stack([signal + np.random.normal(0, 0.8, N),
np.random.normal(0, 1, N),
np.random.normal(0, 1, N)])
y = (signal > 0).astype(int)
X1_tr, X1_te, X2_tr, X2_te, y_tr, y_te = train_test_split(
view1, view2, y, test_size=0.2, random_state=42)
# Scale
sc1 = StandardScaler(); sc2 = StandardScaler()
X1_tr = sc1.fit_transform(X1_tr); X1_te = sc1.transform(X1_te)
X2_tr = sc2.fit_transform(X2_tr); X2_te = sc2.transform(X2_te)
# ── 80 labelled, rest unlabelled ─────────────────────────────
n_labelled = 80
label_idx = np.random.choice(len(y_tr), n_labelled, replace=False)
unlabel_idx = np.setdiff1d(np.arange(len(y_tr)), label_idx)
X1_l, X1_u = X1_tr[label_idx], X1_tr[unlabel_idx]
X2_l, X2_u = X2_tr[label_idx], X2_tr[unlabel_idx]
y_l = y_tr[label_idx]
# ── Co-Training ───────────────────────────────────────────────
def co_train(X1_l, X2_l, y_l, X1_u, X2_u, k=10, max_iter=30):
h1 = LogisticRegression(max_iter=1000)
h2 = LogisticRegression(max_iter=1000)
X1_train, X2_train, y_train = X1_l.copy(), X2_l.copy(), y_l.copy()
X1_pool, X2_pool = X1_u.copy(), X2_u.copy()
for it in range(max_iter):
if len(X1_pool) == 0: break
h1.fit(X1_train, y_train); h2.fit(X2_train, y_train)
# h1 labels top-k for h2
p1 = h1.predict_proba(X1_pool); conf1 = np.max(p1, axis=1)
top1 = np.argsort(conf1)[-k:]
# h2 labels top-k for h1
p2 = h2.predict_proba(X2_pool); conf2 = np.max(p2, axis=1)
top2 = np.argsort(conf2)[-k:]
new_idx = np.union1d(top1, top2)
new_y = np.where(conf1[new_idx] > conf2[new_idx],
np.argmax(p1[new_idx], axis=1),
np.argmax(p2[new_idx], axis=1))
X1_train = np.vstack([X1_train, X1_pool[new_idx]])
X2_train = np.vstack([X2_train, X2_pool[new_idx]])
y_train = np.concatenate([y_train, new_y])
X1_pool = np.delete(X1_pool, new_idx, axis=0)
X2_pool = np.delete(X2_pool, new_idx, axis=0)
return h1, h2
h1, h2 = co_train(X1_l, X2_l, y_l, X1_u.copy(), X2_u.copy())
# ── Evaluate ──────────────────────────────────────────────────
p1 = h1.predict_proba(X1_te)[:,1]
p2 = h2.predict_proba(X2_te)[:,1]
y_co = ((p1 + p2) / 2 > 0.5).astype(int)
acc_co = accuracy_score(y_te, y_co)
# Supervised baseline (view 1 only, 80 labels)
h_sup = LogisticRegression(max_iter=1000).fit(X1_l, y_l)
acc_sup = accuracy_score(y_te, h_sup.predict(X1_te))
print("=== Co-Training — Stock BUY/SELL Classification ===")
print(f"Supervised (80 labels, view1 only): {acc_sup:.4f}")
print(f"Co-Training (80 labels, 2 views): {acc_co:.4f}")
set.seed(42); N <- 2000
# ── Simulate Two-View Data ────────────────────────────────────
signal <- rnorm(N)
view1 <- cbind(signal+rnorm(N,.8), rnorm(N), rnorm(N))
view2 <- cbind(signal+rnorm(N,.8), rnorm(N), rnorm(N))
y <- as.integer(signal > 0)
df1 <- data.frame(view1, y=factor(y))
df2 <- data.frame(view2, y=factor(y))
idx <- sample(1:N, 0.8*N)
train1 <- df1[idx,]; test1 <- df1[-idx,]
train2 <- df2[idx,]; test2 <- df2[-idx,]
# ── 80 labelled only ─────────────────────────────────────────
l_idx <- sample(1:nrow(train1), 80)
u_idx <- setdiff(1:nrow(train1), l_idx)
# ── Simple co-training loop ───────────────────────────────────
X1_l <- as.matrix(train1[l_idx, 1:3]); y_l <- train1$y[l_idx]
X2_l <- as.matrix(train2[l_idx, 1:3])
X1_u <- as.matrix(train1[u_idx, 1:3])
X2_u <- as.matrix(train2[u_idx, 1:3])
X1_te <- as.matrix(test1[,1:3]); y_te <- test1$y
co_train_r <- function(X1_l, X2_l, y_l, X1_u, X2_u, k=10, iters=20) {
library(e1071)
X1t <- X1_l; X2t <- X2_l; yt <- y_l
for(i in 1:iters) {
h1 <- svm(X1t, yt, probability=TRUE, kernel="radial")
h2 <- svm(X2t, yt, probability=TRUE, kernel="radial")
p1 <- attr(predict(h1,X1_u,probability=TRUE),"probabilities")
p2 <- attr(predict(h2,X2_u,probability=TRUE),"probabilities")
c1 <- apply(p1,1,max); c2 <- apply(p2,1,max)
top1 <- order(c1,decreasing=TRUE)[1:k]
top2 <- order(c2,decreasing=TRUE)[1:k]
new <- union(top1,top2)
new_y <- ifelse(c1[new]>c2[new], colnames(p1)[apply(p1[new,,drop=FALSE],1,which.max)],
colnames(p2)[apply(p2[new,,drop=FALSE],1,which.max)])
X1t <- rbind(X1t, X1_u[new,]); X2t <- rbind(X2t, X2_u[new,])
yt <- factor(c(as.character(yt), new_y))
X1_u <- X1_u[-new,]; X2_u <- X2_u[-new,]
if(nrow(X1_u)==0) break
}
list(h1=svm(X1t,yt,probability=TRUE), h2=svm(X2t,yt,probability=TRUE))
}
models <- co_train_r(X1_l,X2_l,y_l,X1_u,X2_u)
p1 <- attr(predict(models$h1,X1_te,probability=TRUE),"probabilities")[,2]
p2 <- attr(predict(models$h2,as.matrix(test2[,1:3]),probability=TRUE),"probabilities")[,2]
y_pred <- factor(as.integer((p1+p2)/2 > 0.5))
acc_co <- mean(y_pred == y_te)
h_sup <- svm(X1_l,y_l,probability=TRUE)
acc_sup <- mean(predict(h_sup,X1_te)==y_te)
cat(sprintf("Supervised Accuracy: %.4f\nCo-Training Accuracy: %.4f\n",acc_sup,acc_co))
The same information as the assumption blocks above, side by side — this is the comparison that decides which method to reach for.
| Algorithm | Assumes | Breaks when |
|---|---|---|
| 3.1 Self-Training | The model's confidence is calibrated | The initial model is poor — confirmation bias compounds its own errors |
| 3.2 Co-Training | Two views, each sufficient alone, conditionally independent given the label | No natural view split exists — the independence premise collapses |