Labels are expensive; raw data is cheap. Semi-supervised learning asks what the unlabelled pool can contribute. The answer depends on an assumption you must be willing to defend: that the unlabelled data comes from the same distribution, and that its geometry says something about the labels. When the assumption holds these methods are close to free; when it fails they make the model actively worse.
By the end of this unit you should be able to:
Use the unlabelled pool as scaffolding to train a classifier that afterwards labels anything.
2 algorithms • Self-Training, Co-Training 3.2Never build a reusable model — propagate labels through a graph over exactly the points you must label.
1 algorithm • Label Propagation| No. | Algorithm | Character | Topic |
|---|---|---|---|
| 3.1 | Self-Training | Pseudo-Labelling • Iterative | 3.1 Inductive Methods |
| 3.2 | Co-Training | Multi-View • Ensemble | 3.1 Inductive Methods |
| 3.3 | Label Propagation | Graph-Based • Manifold | 3.2 Transductive Methods |