Skip to the content

Algorithms on This Page

3.1 Self-Training 3.2 Co-Training
Inductive methods use the unlabelled pool to train a classifier \(f_\theta\) that can afterwards label any new point, including data never seen during training. The unlabelled set is scaffolding; the model is the product.

3.1  Self-Training

Pseudo-LabellingIterative
DEFINITION

Self-Training (Self-Learning) is the simplest semi-supervised method: a classifier trained on labelled data makes predictions on unlabelled data, and the highest-confidence predictions become pseudo-labels added to the training set. The model is retrained iteratively, gradually incorporating more of the unlabelled pool.

3.1.1  Mathematical Foundation

FORMULAE

Algorithm:

  1. Train \(f_\theta\) on labelled set \(\mathcal{L} = \{(\mathbf{x}_i, y_i)\}_{i=1}^l\)
  2. Predict on unlabelled set \(\mathcal{U}\): \(\hat{p}_j = f_\theta(\mathbf{x}_j)\)
  3. Select high-confidence samples: \(\mathcal{U}^* = \{\mathbf{x}_j : \max_c \hat{p}_j(c) \ge \tau\}\)
  4. Add pseudo-labels: \(\mathcal{L} \leftarrow \mathcal{L} \cup \{(\mathbf{x}_j, \hat{y}_j) : \mathbf{x}_j \in \mathcal{U}^*\}\)
  5. Remove from \(\mathcal{U}\); repeat until \(\mathcal{U} = \emptyset\)

Confidence threshold: \(\tau \in [0.8, 0.99]\) — higher reduces noise in pseudo-labels

Risk: Confirmation bias — errors compound if initial model is poor on \(\mathcal{L}\)

3.1.2  How It Works

Self-training rests on confidence calibration — the assumption that a high predicted probability really does track a high chance of being right. (This is distinct from the low-density separation assumption, a geometric claim that the decision boundary passes through sparse regions, which is what justifies transductive SVMs.) Calibration is why confirmation bias is the characteristic failure mode: an over-confident model pseudo-labels its own errors and then trains on them. The choice of threshold \(\tau\) is critical: too low introduces label noise, too high slows progress. Class-balanced selection at each iteration prevents class imbalance propagation. Self-training is particularly effective when the labelled set is small but representative, and the unlabelled set is large and similar in distribution.

3.1.3  Assumptions and Failure Modes

ASSUMES
  • The model's confidence is calibrated
  • Labelled and unlabelled data share a distribution
BREAKS WHEN
  • The initial model is poor — confirmation bias compounds its own errors
  • \(\tau\) is set too low — label noise floods the training set
  • Pseudo-labels are class-imbalanced and the imbalance is left uncorrected

3.1.4  Worked Examples

FINANCE

📁 Credit Scoring with Few Labels

A new bank has only 200 labelled loan accounts (default/no default) but 5,000 unlabelled historical accounts. Self-training expands the effective training set with the pseudo-labels the model is most confident about. The gain is real when the unlabelled pool follows the same distribution as the labelled one — and can be negative when it does not, because early errors are trained on and compound.

IterationLabelledAUC
0 (initial)2000.71
38200.78
101,8000.83
AGRICULTURE

🛰️ Remote Sensing with Few Labels

Only 50 satellite image pixels have been manually labelled as crop type. Self-training with a Random Forest iteratively labels the pool, starting with the pixels most like the 50 labelled by hand. The confidence threshold \(\tau\) is what keeps an early mistake from propagating through the remaining 9,950.

MethodLabelled usedAccuracy
Supervised only5072%
Self-training50+9,50088%
MEDICINE

📋 EHR Classification with Limited Annotations

100 patient records manually annotated for sepsis vs non-sepsis. Self-training on 10,000 unannotated records is attractive because clinical annotation, not data, is the bottleneck. But a model that is over-confident on rare sepsis cases will cheerfully pseudo-label its own misses — so calibration matters more here than headline accuracy.

MethodF1-Score
Supervised (100)0.61
Self-training0.78
Full supervision0.86

3.1.5  Code

Self-Training
import numpy as np
import pandas as pd
from sklearn.semi_supervised import SelfTrainingClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import classification_report, roc_auc_score
from sklearn.model_selection import train_test_split

np.random.seed(42)
N = 5000

# ── Simulate Loan Default Dataset ─────────────────────────────
X_all = np.column_stack([
    np.random.normal(0, 1, N),   # credit score (std)
    np.random.normal(0, 1, N),   # income (std)
    np.random.normal(0, 1, N)    # debt ratio (std)
])
true_y = (X_all[:,0] - X_all[:,2] + 0.5*X_all[:,1] +
           np.random.normal(0, 0.5, N) < 0).astype(int)

# ── Only 200 labelled; rest unlabelled (-1) ───────────────────
X_tr_all, X_te, y_tr_all, y_te = train_test_split(X_all, true_y, test_size=0.2, random_state=42)
n_labelled = 200
y_semi = y_tr_all.copy().astype(float)
unlabelled_idx = np.random.choice(len(y_semi), len(y_semi)-n_labelled, replace=False)
y_semi[unlabelled_idx] = -1  # -1 means unlabelled in sklearn

scaler = StandardScaler()
X_tr_sc = scaler.fit_transform(X_tr_all)
X_te_sc = scaler.transform(X_te)

# ── Supervised baseline (labelled only) ───────────────────────
labelled_mask = y_semi != -1
lr_sup = LogisticRegression(max_iter=1000).fit(X_tr_sc[labelled_mask], y_semi[labelled_mask])
auc_sup = roc_auc_score(y_te, lr_sup.predict_proba(X_te_sc)[:,1])

# ── Self-Training ─────────────────────────────────────────────
base = LogisticRegression(max_iter=1000, random_state=42)
# NOTE: the first argument was renamed `base_estimator` -> `estimator` in
# scikit-learn 1.6. Pass it positionally and the code works on both.
st   = SelfTrainingClassifier(base, threshold=0.85,
                              criterion='threshold', max_iter=20, verbose=False)
st.fit(X_tr_sc, y_semi)
auc_st = roc_auc_score(y_te, st.predict_proba(X_te_sc)[:,1])

print("=== Self-Training — Loan Default ===")
print(f"Supervised (200 labels) AUC : {auc_sup:.4f}")
print(f"Self-Training           AUC : {auc_st:.4f}")
print(f"Iterations used             : {st.n_iter_}")
print(f"Pseudo-labelled samples     : {(st.transduction_ != -1).sum() - n_labelled}")
print("\nClassification Report (Self-Training):")
print(classification_report(y_te, st.predict(X_te_sc), target_names=['No Default','Default']))

# ── Effect of threshold ───────────────────────────────────────
print("\nEffect of confidence threshold:")
for tau in [0.70, 0.80, 0.85, 0.90, 0.95]:
    m = SelfTrainingClassifier(LogisticRegression(max_iter=1000), threshold=tau)
    m.fit(X_tr_sc, y_semi)
    a = roc_auc_score(y_te, m.predict_proba(X_te_sc)[:,1])
    print(f"  τ={tau:.2f}  AUC={a:.4f}  iters={m.n_iter_}")
library(RSSL); set.seed(42)

# ── Simulate Dataset ──────────────────────────────────────────
N <- 2000
X1 <- rnorm(N); X2 <- rnorm(N); X3 <- rnorm(N)
y  <- factor(as.integer(X1 - X3 + 0.5*X2 + rnorm(N,0,.5) < 0))
df <- data.frame(X1,X2,X3,y)

# ── Split: 80% train, 20% test ────────────────────────────────
idx <- sample(1:N, 0.8*N)
tr  <- df[idx,]; te <- df[-idx,]

# ── Simulate only 200 labelled ───────────────────────────────
labelled_idx   <- sample(1:nrow(tr), 200)
unlabelled_idx <- setdiff(1:nrow(tr), labelled_idx)

X_l <- as.matrix(tr[labelled_idx,   1:3])
y_l <- tr$y[labelled_idx]
X_u <- as.matrix(tr[unlabelled_idx, 1:3])
X_te <- as.matrix(te[,1:3])
y_te <- te$y

# ── Supervised baseline ───────────────────────────────────────
sup_model <- glm(y ~ X1+X2+X3, data=tr[labelled_idx,], family=binomial())
p_sup   <- predict(sup_model, te, type="response")
auc_sup <- mean(outer(p_sup[y_te==1], p_sup[y_te==0], ">"))
cat(sprintf("Supervised AUC: %.4f\n", auc_sup))

# ── Self-Training with SVM ────────────────────────────────────
st_model <- SelfLearning(X_l, y_l, X_u,
                          method=LeastSquaresClassifier())
preds_st <- predict(st_model, X_te)
acc_st   <- mean(preds_st == y_te)
cat(sprintf("Self-Training Accuracy: %.4f\n", acc_st))

3.2  Co-Training

Multi-ViewEnsemble
DEFINITION

Co-Training trains two (or more) classifiers on different views (disjoint feature subsets) of the data. Each classifier labels the most confidently-predicted unlabelled examples and passes them to the other classifier for training. The classifiers "teach" each other, leveraging the multi-view redundancy in the data.

3.2.1  Mathematical Foundation

FORMULAE

Two-view assumption: Features can be split into \(\mathbf{x} = (\mathbf{x}^{(1)}, \mathbf{x}^{(2)})\) where each view is sufficient for classification

Algorithm:

  1. Train \(h_1\) on \(\mathcal{L}\) using view 1; train \(h_2\) on \(\mathcal{L}\) using view 2
  2. Each classifier labels top-\(k\) high-confidence unlabelled examples from \(\mathcal{U}\)
  3. Add \(h_1\)'s labels to \(h_2\)'s training set and vice versa
  4. Repeat until \(\mathcal{U}\) is exhausted or convergence

Key assumption: Views are conditionally independent given the class label

Theoretical guarantee (Blum & Mitchell, 1998): if the two views are conditionally independent given the label, and a weak learner on one view alone achieves error slightly better than chance, then unlabelled data can be used to boost the pair to arbitrarily low error. The result is PAC-style learnability, not a per-iteration error bound — in practice the conditional-independence premise rarely holds exactly, and performance degrades gracefully as it is violated.

3.2.2  How It Works

Co-Training requires a natural feature split into two sufficient views. In text classification, view 1 might be content words and view 2 hyperlinks. When natural views don't exist, random feature splits or different algorithms trained on the same features can substitute (Democratic Co-Training). Adding small amounts of confident pseudo-labels per iteration prevents rapid error propagation. Co-Training can outperform self-training when the two views are truly complementary.

3.2.3  Assumptions and Failure Modes

ASSUMES
  • Two views, each sufficient alone, conditionally independent given the label
BREAKS WHEN
  • No natural view split exists — the independence premise collapses
  • The two classifiers agree on everything — they stop teaching each other
  • Too many pseudo-labels are added per round — errors propagate across views

3.2.4  Worked Examples

FINANCE

🏷️ News + Fundamentals Stock Sentiment

View 1: NLP features from news articles; View 2: financial ratios (P/E, EPS growth). Co-training classifies stocks as BUY/HOLD/SELL. The two views are genuinely complementary — news sentiment and fundamentals fail in different ways and on different days — which is the condition co-training needs in order to help at all.

MethodAccuracy
NLP view only65%
Fundamental view61%
Co-Training74%
AGRICULTURE

🔭 Multi-Sensor Crop Classification

View 1: optical satellite bands (RGB, NIR); View 2: SAR radar backscatter. Co-training leverages weather-independent SAR and detail-rich optical data jointly, from only 80 labelled fields. SAR sees through cloud while optical carries texture, so the two views' errors are close to independent — precisely the premise Blum & Mitchell's guarantee rests on.

MethodLabelsAccuracy
Optical only8079%
Co-Training80+4,92092%
MEDICINE

🩻 Medical Image + Clinical Notes

View 1: chest X-ray features (CNN embeddings); View 2: clinical note NLP features. Co-training classifies pneumonia from 50 labelled cases and 2,000 unlabelled records — a setting that matters most where radiologists are scarce. Image features and note features are close to conditionally independent given the diagnosis, which is the assumption the method needs.

MethodSensitivity
X-ray only (50)68%
Co-Training85%

3.2.5  Code

Co-Training
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

np.random.seed(42)
N = 3000

# ── Simulate Two-View Data (Financial: NLP + Fundamentals) ────
# View 1: NLP-like features (3 dims)
# View 2: Fundamental ratios (3 dims)
# True signal split across both views
signal = np.random.normal(0, 1, N)
view1 = np.column_stack([signal + np.random.normal(0, 0.8, N),
                          np.random.normal(0, 1, N),
                          np.random.normal(0, 1, N)])
view2 = np.column_stack([signal + np.random.normal(0, 0.8, N),
                          np.random.normal(0, 1, N),
                          np.random.normal(0, 1, N)])
y = (signal > 0).astype(int)

X1_tr, X1_te, X2_tr, X2_te, y_tr, y_te = train_test_split(
    view1, view2, y, test_size=0.2, random_state=42)

# Scale
sc1 = StandardScaler(); sc2 = StandardScaler()
X1_tr = sc1.fit_transform(X1_tr); X1_te = sc1.transform(X1_te)
X2_tr = sc2.fit_transform(X2_tr); X2_te = sc2.transform(X2_te)

# ── 80 labelled, rest unlabelled ─────────────────────────────
n_labelled = 80
label_idx = np.random.choice(len(y_tr), n_labelled, replace=False)
unlabel_idx = np.setdiff1d(np.arange(len(y_tr)), label_idx)

X1_l, X1_u = X1_tr[label_idx], X1_tr[unlabel_idx]
X2_l, X2_u = X2_tr[label_idx], X2_tr[unlabel_idx]
y_l = y_tr[label_idx]

# ── Co-Training ───────────────────────────────────────────────
def co_train(X1_l, X2_l, y_l, X1_u, X2_u, k=10, max_iter=30):
    h1 = LogisticRegression(max_iter=1000)
    h2 = LogisticRegression(max_iter=1000)
    X1_train, X2_train, y_train = X1_l.copy(), X2_l.copy(), y_l.copy()
    X1_pool, X2_pool = X1_u.copy(), X2_u.copy()
    
    for it in range(max_iter):
        if len(X1_pool) == 0: break
        h1.fit(X1_train, y_train); h2.fit(X2_train, y_train)
        # h1 labels top-k for h2
        p1 = h1.predict_proba(X1_pool); conf1 = np.max(p1, axis=1)
        top1 = np.argsort(conf1)[-k:]
        # h2 labels top-k for h1
        p2 = h2.predict_proba(X2_pool); conf2 = np.max(p2, axis=1)
        top2 = np.argsort(conf2)[-k:]
        
        new_idx = np.union1d(top1, top2)
        new_y   = np.where(conf1[new_idx] > conf2[new_idx],
                           np.argmax(p1[new_idx], axis=1),
                           np.argmax(p2[new_idx], axis=1))
        X1_train = np.vstack([X1_train, X1_pool[new_idx]])
        X2_train = np.vstack([X2_train, X2_pool[new_idx]])
        y_train  = np.concatenate([y_train, new_y])
        X1_pool  = np.delete(X1_pool, new_idx, axis=0)
        X2_pool  = np.delete(X2_pool, new_idx, axis=0)
    return h1, h2

h1, h2 = co_train(X1_l, X2_l, y_l, X1_u.copy(), X2_u.copy())

# ── Evaluate ──────────────────────────────────────────────────
p1 = h1.predict_proba(X1_te)[:,1]
p2 = h2.predict_proba(X2_te)[:,1]
y_co = ((p1 + p2) / 2 > 0.5).astype(int)
acc_co = accuracy_score(y_te, y_co)

# Supervised baseline (view 1 only, 80 labels)
h_sup = LogisticRegression(max_iter=1000).fit(X1_l, y_l)
acc_sup = accuracy_score(y_te, h_sup.predict(X1_te))

print("=== Co-Training — Stock BUY/SELL Classification ===")
print(f"Supervised (80 labels, view1 only): {acc_sup:.4f}")
print(f"Co-Training (80 labels, 2 views):   {acc_co:.4f}")
set.seed(42); N <- 2000

# ── Simulate Two-View Data ────────────────────────────────────
signal <- rnorm(N)
view1  <- cbind(signal+rnorm(N,.8), rnorm(N), rnorm(N))
view2  <- cbind(signal+rnorm(N,.8), rnorm(N), rnorm(N))
y      <- as.integer(signal > 0)
df1    <- data.frame(view1, y=factor(y))
df2    <- data.frame(view2, y=factor(y))

idx    <- sample(1:N, 0.8*N)
train1 <- df1[idx,]; test1 <- df1[-idx,]
train2 <- df2[idx,]; test2 <- df2[-idx,]

# ── 80 labelled only ─────────────────────────────────────────
l_idx  <- sample(1:nrow(train1), 80)
u_idx  <- setdiff(1:nrow(train1), l_idx)

# ── Simple co-training loop ───────────────────────────────────
X1_l <- as.matrix(train1[l_idx, 1:3]); y_l <- train1$y[l_idx]
X2_l <- as.matrix(train2[l_idx, 1:3])
X1_u <- as.matrix(train1[u_idx, 1:3])
X2_u <- as.matrix(train2[u_idx, 1:3])
X1_te <- as.matrix(test1[,1:3]); y_te <- test1$y

co_train_r <- function(X1_l, X2_l, y_l, X1_u, X2_u, k=10, iters=20) {
  library(e1071)
  X1t <- X1_l; X2t <- X2_l; yt <- y_l
  for(i in 1:iters) {
    h1 <- svm(X1t, yt, probability=TRUE, kernel="radial")
    h2 <- svm(X2t, yt, probability=TRUE, kernel="radial")
    p1 <- attr(predict(h1,X1_u,probability=TRUE),"probabilities")
    p2 <- attr(predict(h2,X2_u,probability=TRUE),"probabilities")
    c1 <- apply(p1,1,max); c2 <- apply(p2,1,max)
    top1 <- order(c1,decreasing=TRUE)[1:k]
    top2 <- order(c2,decreasing=TRUE)[1:k]
    new  <- union(top1,top2)
    new_y <- ifelse(c1[new]>c2[new], colnames(p1)[apply(p1[new,,drop=FALSE],1,which.max)],
                    colnames(p2)[apply(p2[new,,drop=FALSE],1,which.max)])
    X1t  <- rbind(X1t, X1_u[new,]); X2t <- rbind(X2t, X2_u[new,])
    yt   <- factor(c(as.character(yt), new_y))
    X1_u <- X1_u[-new,]; X2_u <- X2_u[-new,]
    if(nrow(X1_u)==0) break
  }
  list(h1=svm(X1t,yt,probability=TRUE), h2=svm(X2t,yt,probability=TRUE))
}

models <- co_train_r(X1_l,X2_l,y_l,X1_u,X2_u)
p1 <- attr(predict(models$h1,X1_te,probability=TRUE),"probabilities")[,2]
p2 <- attr(predict(models$h2,as.matrix(test2[,1:3]),probability=TRUE),"probabilities")[,2]
y_pred <- factor(as.integer((p1+p2)/2 > 0.5))
acc_co <- mean(y_pred == y_te)

h_sup <- svm(X1_l,y_l,probability=TRUE)
acc_sup <- mean(predict(h_sup,X1_te)==y_te)
cat(sprintf("Supervised Accuracy: %.4f\nCo-Training Accuracy: %.4f\n",acc_sup,acc_co))

At a Glance

The same information as the assumption blocks above, side by side — this is the comparison that decides which method to reach for.

AlgorithmAssumesBreaks when
3.1 Self-TrainingThe model's confidence is calibratedThe initial model is poor — confirmation bias compounds its own errors
3.2 Co-TrainingTwo views, each sufficient alone, conditionally independent given the labelNo natural view split exists — the independence premise collapses