Self-Supervised Learning
Self-supervised learning creates training targets from the data itself rather than requiring a human label for each example. Tasks include predicting hidden text and learning representations from different views of an image.
A model sees “The cat sat on the [MASK]” and learns to predict the hidden word from the original sentence. No annotator needs to write that target: the corpus supplies it. BERT uses masked-token prediction; a next-token language model instead uses the following token as its target.
For images, a contrastive method can train representations of two augmented views of one photograph to agree while distinguishing other photographs. The views supply a relationship to learn without a manual class label. These are different objectives within the same broad approach.
Self-supervision can make large unlabeled collections useful for representation learning. It is often grouped under unsupervised learning, although the optimization has generated targets. Success on the pretraining objective does not establish truthfulness or task suitability. Test the representation or adapted model on the task it will actually serve.
Sources
- Devlin et al.: BERT — Describes masked-language-model pretraining from unlabeled text.
- Chen et al.: SimCLR — Learns visual representations with contrastive objectives and augmented views rather than class labels.
Go deeper
- Keras: Contrastive pretraining with SimCLR docs
Follow image augmentations, contrastive pretraining and evaluation with a labeled classifier.