Weak supervision
Weak supervision trains models using imperfect supervision, such as noisy labels, incomplete labels or coarse labels. Programmatic weak supervision combines rules and other approximate label sources to create training targets.
Also known as: weakly supervised learning
A fraud team could label transactions using business rules and vendor flags. One rule flags unusually large purchases; another flags unfamiliar devices. These sources can conflict or miss legitimate exceptions. In a Snorkel-style pipeline, a label model combines their outputs into probabilistic targets, then a separate classifier learns from those targets.
This distinction matters: the classifier is not simply applying the rules at prediction time. It learns from their approximate supervision and may generalize beyond their coverage. Its quality still depends on what the rules miss and how their errors relate.
Evaluate on independently reviewed labels. Several rules copied from the same assumption can agree and still be wrong. Unlike unsupervised learning, the model receives task targets; unlike self-supervised learning, the targets in this example come from external labeling rules.
Sources
- Ratner et al.: Snorkel — Describes labeling functions with unknown accuracy and correlations, denoising and probabilistic training labels.
- Zhou: A Brief Introduction to Weakly Supervised Learning — Distinguishes incomplete, inexact and inaccurate supervision.
Go deeper
- Snorkel: API documentation docs
Inspect labeling functions, label models and utilities for combining programmatic supervision.