Semi-supervised learning
Semi-supervised learning trains a model using both labeled and unlabeled examples. The unlabeled data contributes structure or additional training signals to help with the task defined by the labeled data.
A support inbox could contain 200 human-labeled urgency examples and thousands of unlabeled tickets. A self-training method first fits a classifier to the labeled tickets, predicts labels for the unlabeled ones and adds selected confident predictions to training. Another approach propagates labels through a graph of similar examples.
The extra examples help only if the assumptions are useful. Confident errors can become new training targets, and unrelated unlabeled messages can distort the learned boundary. Compare against the labeled-only baseline on a separate, trusted evaluation set; confidence alone does not establish a correct label.
Unsupervised learning has no supplied task labels, while semi-supervised learning has some. Self-supervised learning creates targets from inputs; its representations can be used within a semi-supervised pipeline. State how the unlabeled examples enter training before comparing results.
Sources
- scikit-learn: Semi-supervised learning — Documents self-training with predicted labels, calibration concerns and graph-based label propagation.
Go deeper
- scikit-learn: Label propagation on circles docs
See how a few supplied labels spread through a graph of unlabeled examples.