Skip to the content

What This Unit Is About

Labels are expensive; raw data is cheap. Semi-supervised learning asks what the unlabelled pool can contribute. The answer depends on an assumption you must be willing to defend: that the unlabelled data comes from the same distribution, and that its geometry says something about the labels. When the assumption holds these methods are close to free; when it fails they make the model actively worse.

Learning Outcomes

By the end of this unit you should be able to:

  1. Distinguish inductive from transductive methods, and pick the right one for whether you must serve predictions later.
  2. Explain confirmation bias in self-training, and how the confidence threshold controls it.
  3. State the two-view assumption co-training depends on, and recognise when a dataset does not satisfy it.
  4. Build a similarity graph and propagate labels across it, with hard and with soft clamping.
  5. Judge when a semi-supervised method will help — and when it will amplify the errors of a weak initial model.

Topics in This Unit

3.1

🔁 Inductive Methods

Use the unlabelled pool as scaffolding to train a classifier that afterwards labels anything.

2 algorithms • Self-Training, Co-Training
3.2

🕸️ Transductive Methods

Never build a reusable model — propagate labels through a graph over exactly the points you must label.

1 algorithm • Label Propagation

Every Algorithm in Semi-Supervised Learning

No.AlgorithmCharacterTopic
3.1Self-TrainingPseudo-Labelling • Iterative3.1 Inductive Methods
3.2Co-TrainingMulti-View • Ensemble3.1 Inductive Methods
3.3Label PropagationGraph-Based • Manifold3.2 Transductive Methods