Statistical classification
Statistical classification uses a model learned from data to assign an input to a category. A classifier may produce scores or probabilities that a decision rule converts into labels.
A fraud team could train on transactions labeled fraudulent or legitimate. A model produces a fraud probability, and a chosen threshold sends high-scoring transactions to human review. The predicted label and the decision to block a payment are separate steps; the business process determines what follows the score.
For an illustrative set with 10 fraudulent transactions and 990 legitimate ones, labeling every transaction legitimate gives 99% accuracy while catching no fraud. Precision and recall show different costs: how many flagged transactions are fraudulent, and how much of the fraud is caught. Changing the threshold changes that tradeoff.
Regression in ML usually predicts numbers, while classification predicts categories. Logistic regression is a naming exception: it models class probabilities and is commonly used as a classifier. Text generation can use token probabilities internally, but generating an open-ended response is a different application from choosing a fixed task label.
Sources
- Google: Classification — Introduces classification scores, logistic regression and decision thresholds.
- Google: Accuracy, recall and precision — Defines classification metrics and explains why accuracy can mislead on imbalanced data.
Go deeper
- scikit-learn: Model evaluation docs
Compare confusion matrices, precision, recall and score-based metrics for a classifier.